scieee Open visual document viewer

Construction of a map-based reference genome sequence for barley, Hordeum vulgare L.

Beier, S.,Himmelbach, A.,Colmsee, Chr.,Zhang, X.-Q.,Barrero, R. A.,Zhang, Q.,Li, L.,Bayer, M.,Bolser, D.,Taudien, St.,Groth, M.,Felder, M.,Hastie, A.,Simkova, H.,Stankova, H.,Vrana, J.,Chan, S.,Munoz-Amatriain, M.,Ounit, R.,Wanamaker, St.,Schmutzer, Th.,

Full text

Da a Desc ip o : Cons uc ion o a map-based e e ence genome sequence o ba ley, Ho deum ulga e L. Sebas ian Beie e al. # Ba ley (Ho deum ulga e L.) is a ce eal g ass mainly used as animal odde and aw ma e ial o he mal ing indus y. The map-based e e ence genome sequence o ba ley c . ‘Mo ex’was cons uc ed by he In e na ional Ba ley Genome Sequencing Conso ium (IBSC) using hie a chical sho gun sequencing. He e, we epo he expe imen al and compu a ional p ocedu es o (i) sequence and assemble mo e han 80,000 bac e ial a ificial ch omosome (BAC) clones along he minimum iling pa h o a genome-wide physical map, (ii) find and alida e o e laps be ween adjacen BACs, (iii) cons uc 4,265 non- edundan sequence sca olds ep esen ing clus e s o o e lapping BACs, and (i ) o de and o ien hese BAC clus e s along he se en ba ley ch omosomes using posi ional in o ma ion p o ided by dense gene ic maps, an op ical map and ch omosome con o ma ion cap u e sequencing (Hi-C). In eg a i e access o hese sequence and mapping esou ces is p o ided by he ba ley genome explo e (BARLEX). Design Type(s) genome assembly Measu emen Type(s) whole genome sequencing assay Technology Type(s) DNA sequencing Fac o Type(s) lib a y p epa a ion Sample Cha ac e is ic(s) Ho deum ulga e Co espondence and eques s o ma e ials should be add essed o M.M. (email: masche @ipk-ga e sleben.de). #A ull lis o au ho s and hei a filia ions appea s a he end o he pape . OPEN Recei ed: 26 Augus 2016 Accep ed: 9Feb ua y 2017 Published: 27 Ap il 2017 www.na u e.com/scien i icda a SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 1 Backg ound & Summa y Ba ley (Ho deum ulga e L.) is a ce eal g ass o g ea ag onomical impo ance. The goal o he In e na ional Ba ley Genome Sequencing Conso ium (IBSC) is he cons uc ion o a map-based e e ence sequence assembly o ba ley cul i a ‘Mo ex’by means o hie a chical sho gun sequencing 1 . Towa ds his aim, he ba ley genomics communi y has de eloped an a ay o genome-wide physical and gene ic mapping esou ces. These include lib a ies o bac e ial a ificial ch omosomes (BACs) 2 , a genome-wide physical map 3 , a d a whole genome sho gun (WGS) assembly 4 and an ul a-dense gene ic map 5 . The las s age on he oad owa ds he e e ence genome is he sho gun sequencing o BAC clones along a minimum iling pa h o he genome defined by he physical map. The ad ances in high- h oughpu sequencing echnology enabled his ask o be comple ed in a much sho e ime ame han was equi ed o he comple ion o , o ins ance, he human 6 and maize 7 genomes. In addi ion o he gene a ion o BAC aw sequence da a, we cons uc ed (i) physical genome maps by single-molecule op ical mapping in nanochannels 8 and by ch omosome con o ma ion cap u e sequencing (Hi-C) 9,10 , and (ii) a high- esolu ion gene ic map o a la ge bi-pa en al mapping popula ion h ough geno yping-by- sequencing 11 . We unde ook he sequence assembly o indi idual BACs, he cons uc ion o la ge sequence sca olds by me ging sequences om adjacen clones and he in eg a ion o hese supe - sca olds wi h he a ious genome-wide mapping esou ces cons uc ed in he p esen e o as well as hose published p e iously 3,5 . The final ou come o his app oach was he cons uc ion o ‘pseudomolecules’, i.e., con iguous sequence sca olds ep esen ing he se en ch omosomes o ba ley. We ha e submi ed he ele an aw da a o public sequence da a a chi es, made analysis esul s a ailable unde pe manen digi al objec iden ifie s (DOIs) and en e ed he posi ional in o ma ion used o pseudomolecule cons uc ion in o a bespoke in o ma ion managemen sys em, he BARLEX genome explo e 12 . He e, we gi e (i) a comp ehensi e o e iew o da ase s used o assembling he ba ley genome and me hods employed in hei gene a ion, (ii) a de ailed desc ip ion o we -lab p ocedu es o BAC sequencing and he bioin o ma ics wo kflow o he sequence assembly and da a in eg a ion p ocedu es oge he wi h an ou line o (iii) hei b owsable p esen a ion in an online da abase. These esou ces documen he cons uc ion o he map-based e e ence sequence o he ba ley genome and will enable esea che s o inspec he e idence used o assemble, o de and o ien sequence sca olds and may guide he u he imp o emen o he genome sequence wi h complemen a y da a se s. Me hods The main s eps o he cons uc ion o he map-based e e ence sequence o he ba ley genome we e (i) sho gun and ma e-pai sequencing o BAC clones, (ii) sequence assembly o indi idual BAC clones and (iii) he cons uc ion o a pseudomolecule sequences by me ging he sequences o adjacen BACs in o supe -sca olds and o de ing hese using a ious sou ces o posi ional in o ma ion such as physical maps, op ical map and ch omosome con o ma ion cap u e. A schema ic o e iew o ou expe imen al p ocedu es is gi en in Fig. 1. BAC sequencing Iden ifica ion and analysis o gene-con aining BACs. Isola ion o gene-con aining BACs, cons uc ion o a minimal iling pa h (MTP), sequencing o MTP clones and he anno a ion o genes we e essen ially as desc ibed p e iously 13 . Sho gun and ma e-pai sequencing o MTP-BACs. Sequencing o MTP-BACs was conduc ed in ou labo a o ies (Leibniz Ins i u e on Aging—F i z Lipmann Ins i u e (FLI) Jena, Leibniz Ins i u e o Plan Gene ics and C op Plan Resea ch (IPK) Ga e sleben, Beijing Genomics Ins i u e (BGI) and Ea lham Ins i u e (EI) No wich). Depending on he ins umen a ion and es ablished p o ocols, cus omized app oaches we e aken o sequence he ba ley MTP BACs. Ba ley ch omosomes 1H, 3H and 4H (IPK and FLI) Sho gun sequencing o MTP BACs Du ing he ini ial phase, BACs mos ly om ch omosome 3H (4870 clones) and a small numbe o clones om o he ch omosomes (34 om 1H; 31 om 2H; 50 om 4H; 101 om 5H; 33 om 6H; 64 om 7H; 107 om ‘0H’) we e sho gun sequenced using he Roche/454 GS FLX de ice (Da a Ci a ion 1, Da a Ci a ion 2, Da a Ci a ion 3, Da a Ci a ion 4, Da a Ci a ion 5, Da a Ci a ion 6, Da a Ci a ion 7, Da a Ci a ion 8, Da a Ci a ion 9). BAC DNA was p epa ed using a modified alkaline lysis p o ocol 14 . Cons uc ion o ba coded 454 sequencing lib a ies and sequencing using he Roche pla o m we e pe o med as desc ibed 15,16 . The emaining BAC clones om ch omosomes 1H, 3H and 4H we e sho gun sequenced employing Illumina ins umen s. BAC DNA isola ion, lib a y cons uc ion, sequencing-by-syn hesis (pai ed-end, 2 × 100 cycles) using he Illumina HiSeq2000 de ice was pe o med as desc ibed 17 (Da a Ci a ion 10, Da a Ci a ion 11, Da a Ci a ion 12, Da a Ci a ion 13). Pools o up o 667 BACs we e indi idually ba coded and sequenced on one HiSeq2000 lane. In addi ion, he Illumina GAIIx, HiSeq2500 and MiSeq machines we e u ilized o sequence pools o up o 384 clones pe lane as desc ibed p e iously 17 . www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 2 Ma e-pai sequencing o MTP BACs Fo sca olding o ch omosomes 1H, 3H and 4H s anda d Illumina Nex e a ma e-pai lib a ies (span size: 8 kb) o BAC pools up o 384 BACs we e cons uc ed and sequenced using he Illumina HiSeq2000 (pai ed end, 2 × 100 cycles) and MiSeq (pai ed end, 2 × 250 cycles) as desc ibed 17 (Da a Ci a ion 14, Da a Ci a ion 15). Ba ley ch omosomes 5H, 6H and 7H (BGI) Sho gun sequencing o MTP BACs Bac e ial s a e cul u es we e inocula ed in 0.4 ml 2 × YT liquid medium 18 supplemen ed wi h chlo amphenicol (17.5 μgml 1 ) in 2 ml polyp opylene 96-deep well-pla es sealed wi h gas-pe meable oil and incuba ed a 37 °C o 14 h in a shaking incuba o (210 .p.m.). Fo DNA isola ion duplica es o cul u es (1 ml 2 × YT liquid medium con aining 17.5 μgml 1 chlo amphenicol) we e inocula ed wi h 50 μl s a e cul u e and incuba ed (37 °C, 14 h, 210 .p.m.). BAC DNA was isola ed using he alkaline lysis me hod essen ially as desc ibed p e iously 17 . The DNA was dissol ed (o e nigh , 4 °C) in 64 μlTE (pH 8.0) con aining RNase A (30 μgml 1 ) and s o ed a −20 °C. BAC plasmid DNA (0.5–2.0 μgin60μl) was andomly agmen ed by ocused-ul asonica o (Co a is LE220 ins umen : 21% du y ac o , 500 PIP, 500 cycles pe bu s , 70 s ea men ime) in 96-well pla es (Axygen, PCR-96M2-HS-C) o an a e age size o 250–750 bp. The DNA agmen s we e pu ified using magne ic beads (GeneOn Pu ifica ion ki , GO-PCRC-5000) acco ding o he manu ac u e ’s ins uc ions. DNA was p ecipi a ed by adding 10 μl magne ic bead suspension and 75 μl Binding Bu e . The samples we e mixed and incuba ed a oom empe a u e o 5 min. Beads con aining he DNA we e eclaimed by using a magne (96S Supe Magne Pla e, ALPAQUA, A001322), and he clea supe na an was disca ded. The beads we e washed wice wi h 200 μl o 70% e hanol and d ied comple ely. Fo he elu ion o DNA he beads we e suspended in 42 μl Elu ion Bu e (EB, 10 mM T is-Cl, pH 8.5) and incuba ed (5 min). The pla e was placed on he magne , and he supe na an (40 μl) was ans e ed in o new 96-well pla es. End- epai and A-Tailing we e pe o med as desc ibed 19 . The eac ion clean-ups we e pe o med wi h GeneOn magne ic beads as desc ibed abo e. Ba code adap e s (1 μl, 20 μM) o he fi s index we e liga ed o he s icky ends o DNA agmen s by using T4 DNA ligase 19 , incuba ed a 16 °C o a leas 12 h. Each indi idual sample was p o ided wi h a di e en ba code o a se o 384 di e en indices (adap e and ba code sequences a e a ailable upon eques ). Equal olumes o he 384 indi idually ba coded adap e -liga ed p oduc s we e pooled. The pooled DNA was p ecipi a ed by adding 20 μl GeneOn magne ic beads and 650 μl Binding Bu e (GeneOn Pu ifica ion Ki , GO-PCRC-5000) BAC DNA p epa a ion Pai ed-end lib a y cons uc ion Ma e-pai lib a y cons uc ion Quan i ica ion, pooling and size ac iona ion Sequencing-by- syn hesis (Illumina) Remo al o low quali y sequences & con amina ion Indi idual BAC assembly Quan i ica ion, pooling and size ac iona ion Sequencing-by- syn hesis (Illumina) Remo al o low quali y sequences & con amina ion Mapping & indi idual BAC sca olding Indi idual BAC sca olds BAC sca olds FPC / BES da a POPSEQ map Bionano map da a Con o ma ion cap u e da a (HiC / TCC) BAC o e lap clus e s BLAST analysis Non- edundan sequence Con o ma ion cap u e map (HiC map) AGP gene a ion & Pseudomolecule sequence Figu e 1. Assembly wo kflow. (a) Assembly o indi idual BAC clones om pai ed-end and ma e-pai ead da a. (b) Da a in eg a ion p ocedu es o pseudomolecule cons uc ion. www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 3 o 500 μl pooled DNA. The suspension was mixed and incuba ed a oom empe a u e o 5 min. The beads con aining he DNA we e eclaimed using a magne , and he clea supe na an was disca ded. The beads we e washed wice wi h 500 μl o 70% e hanol and d ied comple ely. The DNA was elu ed in 52 μl EB. The sample was size-sepa a ed by using s anda d aga ose gel elec opho esis (2% aga ose gel, HyAga ose, 16250). DNA was e ealed using e hidium b omide and exci a ion by isible blue ligh emi ed om a Da k Reade blue ligh ansillumina o (Cla e Chemical Resea ch) o selec he a ge agmen s (580–620 bp). The a ge egion was ex ac ed in 27 μl EB using he QIAquick Gel Ex ac ion ki (QIAGEN). The second index was in oduced using he adap e -liga ed p oduc s as empla e DNA (98 °C o 30 s, 10 cycles o : 98 °C o 10 s, 65 °C o 30 s and 72 °C o 30 s, final ex ension 72 °C o 5 min) (Enzyma ics, CM0075) and PCR p oduc s ( a ge egion: 580–620 bp) we e eco e ed by aga ose gel elec opho esis (2% aga ose gel, HyAga ose, 16250) as desc ibed abo e. Index p ime s we e used o ba coding each 384 pooled BAC samples (index p ime sequences a e a ailable upon eques ). The a e age size o he PCR p oduc s was de e mined by using an Agilen 2100 Bioanalyze (Agilen DNA 1,000 Reagen s). Typical a e age size o he lib a ies was be ween 574 o 674 bp. PCR p oduc s we e quan ified using eal- ime PCR and pooled o sequencing in equal p opo ion 20 . Pai ed-end sequencing (2 × 100 cycles; fi s index: 11 cycles, second index: 8 cycles) was pe o med on he Illumina HiSeq2000 pla o m (Da a Ci a ion 16, Da a Ci a ion 17, Da a Ci a ion 18). Ma e-pai sequencing o MTP BACs Fo he cons uc ion o ma e-pai lib a ies (10 and 20 kb span size), 96 BACs co esponding o 6 μg DNA we e pooled in o one ube. The DNA was agmen ed o 10 o 20 kb by using he Hyd oShea DNA Shea ing sys em om GeneMachines (10 kb: la ge assembly, speed code 12, cycles 12, olume 250 μl; 20 kb: la ge assembly, speed code 13, cycles 20, olume 150 μl). Following DNA agmen a ion, he agmen s we e pu ified by using 0.6 olumes magne ic beads (Axygen, MAG-PCR-CL-250). The samples we e mixed and incuba ed a oom empe a u e o 10 min. Beads con aining he DNA we e eclaimed by using a magne pla e (96S Supe Magne Pla e, ALPAQUA, A001322), and he clea supe na an was disca ded. The beads we e washed wice wi h 500 μl o 70% e hanol and d ied comple ely. Fo he elu ion o DNA he beads we e esuspended in 80 μl EB. End- epai and bio in-labeling we e pe o med as desc ibed 21 . End- epai ed DNA was pu ified using 0.6 olumes magne ic beads (Axygen, MAG-PCR- CL-250) as desc ibed o he pu ifica ion o hyd o-shea ed DNA. The DNA was elu ed in 79 μl EB. 20 kb lib a ies (20–26 kb ange) we e size-selec ed using aga ose gel (0.6%) elec opho esis. The liga ion o he lib a ies, was pe o med by adding 1 μl Ba code Adap o (20 μM, sequences a e a ailable upon eques ), 10 μl T4 DNA ligase (Enzyma ics, L603-HC) in a o al olume o 100 μl (20 °C, 15 min). 15 indi idually ba coded adap o -liga ed DNAs (10 kb) we e pooled in equimola manne and size- ac iona ed (9–11 kb) using aga ose gel (0.6%) elec opho esis. DNA ci cula iza ion and emo al o non-ci cula ized DNA was as desc ibed 21 . The DNA was isola ed om he gel using he QIAquick Gel Ex ac ion ki as desc ibed by he manu ac u e (QIAGEN). Ci cula DNA was agmen ed using he Co a is S2 de ice (10% du y cycle, 10 in ensi y, 1,000 bu s s pe second, 22 min (11 min) ea men ime o 10 kb (20 kb) lib a ies in TC13 Co a is ubes), and bio inyla ed agmen s de i ed om ue ma e-pai liga ion e en s we e pu ified using s ep a idin-coupled Dynabeads (M-280, In i ogen) 19 . Ends o he DNA agmen s we e epai ed and p o ided wi h Illumina pai ed-end adap e s as desc ibed o he cons uc ion o sho gun lib a ies. The bead-bound DNA was PCR-amplified using Phusion polyme ase (NEB) (98 °C o 30 s, 18 cycles o : 98 °C o 10 s, 65 °C o 30 s, 72 °C o 30 s and a final ex ension: 72 °C o 5 min) using manu ac u e ’s p o ocols (NEB). Size-selec ion was essen ially pe o med as desc ibed o sho gun lib a y cons uc ion. Fo he 10 kb (20 kb) ma e-pai lib a ies, DNA in he size ange be ween 270–420 bp (400–600 bp) was isola ed and pu ified using he QIAquick Gel Ex ac ion ki acco ding o manu ac u e ’s ins uc ions (QIAGEN). The a e age size o he pai ed-end BAC lib a ies was de e mined elec opho e ically using an Agilen 2100 Bioanalyze (Agilen DNA 1,000 Reagen s). Lib a ies we e quan ified using Real-Time PCR 20 . The ma e-pai lib a ies we e pai ed-end sequenced using he Illumina HiSeq2500 de ice (10 kb lib a y: 150 cycles, 20 kb ma e-pai lib a y 50 cycles). Raw da a a e a ailable as Da a Ci a ion 19, Da a Ci a ion 20, Da a Ci a ion 21). Ba ley ch omosomes 2H and 0H (EI) Sho gun sequencing o MTP BACs QRep 384 Pin Replica o s (Molecula De ices, New Mol on, UK) we e used o inocula e clones om s ock pla es in o 384 squa e deep well cul u e pla es con aining 140 μl 2 × YT media supplemen ed wi h 12.5 μgml 1 chlo amphenicol 18 . The cul u e pla es we e sealed wi h a gas pe meable seal and incuba ed o 22 h a 37 °C in a shaking incuba o (200 .p.m.). Cells we e ha es ed by cen i uga ion (20 min, 3,220 g, 4 °C), he supe na an was disca ded. BAC DNA was p epa ed using a modified alkaline lysis p o ocol (Beckman Coul ie , High Wycombe, UK). Cell pelle s we e esuspended in 8 μl o Resuspension Bu e (RE1) using a Mic opla e Shake TiMix 5 con ol (Edmund-Buehle , Hechingen, Ge many) (10 min, 1,400 .p.m.). Cells we e lysed by adding 8 μl o he lysis solu ion (L2). A e shaking (5 min, 500 .p.m.) 8 μl o cold Neu alisa ion Bu e (N3) we e added. The pla e was shaken (10 min, 500 .p.m.) ollowed by a cen i uga ion (20 min, 3,220 g, 4 °C). The clea supe na an (14.33 μl) was ans e ed www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 4 o a 384 well PCR pla e, which con ained 1 μl o CosMc beads pe well. The pla e was mixed b iefly (500 .p.m.), 10 μl o isop opanol was added and he suspension was mixed b iefly again (500 .p.m.). The pla e was incuba ed a oom empe a u e o 15 min o allow p ecipi a ion o he DNA on o he beads. The pla e con aining he DNA p ecipi a e was mo ed on o a 96 pin 384 well pla e compa ible magne (Alpaqua, Be e ley, MA, USA) and le o 5 min o he beads o pelle . The supe na an was disca ded and he beads we e washed h ee imes wi h 20 μl 70% e hanol while placed in he magne and ai d ied ( oom empe a u e, 5 min). The DNA was elu ed om he beads in 20 μl o 10 mM T is HCl (pH 8.0) and ans e ed o a esh 384 well PCR pla e. To emo e con amina ing hos E. coli gDNA samples we e ea ed wi h Epicen e Plasmid Sa e ATP dependen DNase (Cambio, Camb idge, UK), which diges s he agmen ed E. coli and nicked BAC DNA bu lea es supe coiled BAC DNA in ac . To 20 μl o DNA 2.5 μl o 10x Reac ion bu e , 1 μl 25 mM ATP, 0.1 μl ATP dependen DNase (10 u μl 1 ) and 1.4 μl wa e was added, and he samples we e incuba ed a 37 °C (8 h) ollowed by 70 °C (20 min) o inac i a e he DNase. Sequencing lib a ies (single index) om he ini ial six een 384 well pla es o BACs (2H ch omosome) we e cons uc ed in 384 well PCR pla es (Fo i ude, Wo on, UK) using he Epicen e Nex e a Ki (Epicen e, Madison, WI, USA) and Robus 2G Taq polyme ase (Kapa Biosciences, London, UK). The 384 adap e oligos wi h 9 bp ba codes each wi h a hamming dis ance o 4 (adap e sequences a e a ailable upon eques ) we e designed using s anda d guidelines 22 . B iefly, 1 μl o BAC DNA, 1 μl Nex e a HMW 5 × Reac ion Bu e , 1 μl o Nex e a Enzyme (dilu ed 50- old in 50% glyce ol, 0.5 × TE pH 8.0) and 2 μlo wa e we e combined and incuba ed (5 min, 55 °C) as desc ibed 23 . Fo he dena u a ion o he Tn5 polyme ase, 15 μl PB Bu e (Qiagen, Manches e , UK) and o he eac ion clean-up, 20 μl AMPu e XP (Beckman, High Wycombe, UK) beads we e added using a Calipe Sciclone Robo (Pe kin Elme , Co en y, UK). Following an incuba ion (5 min, oom empe a u e), he p ecipi a ed agmen ed DNA was pu ified using a 96 well ing Magne (Alpaqua, Be e ly, MA, USA). The beads we e washed wice wi h 20 μl 70% e hanol while placed in he magne be o e being ai d ied o 5 min. The agmen ed DNA was elu ed in 5 μl 10 mM T is HCl, pH 8.0 and ans e ed o a esh 384 well PCR pla e. To 5 μl pu ified, agmen ed DNA 2 μl o 5 × 2G B Reac ion bu e , 0.2 μl o 10 mM dNTPs, 0.1 μl o Robus 2G Taq polyme ase, 0.2 μl o 50 × Nex e a P ime Cock ail and 2.5 μl 0.2 μM ba coded P2 adap e p ime we e added in a o al eac ion olume o 10 μl and amplified acco ding o he ollowing he mal cycling p ofile: 72 °C o 3 min, 95 °C o 1 min, ollowed by 21 cycles o 95 °C o 10 s, 65 °C o 20 s and 72 °C o 3 min. Pos amplifica ion he DNA concen a ion was de e mined using he Quan -I Picog een dsDNA assay (The mo Fishe , Camb idge, UK). Lib a y DNA concen a ions ypically anged om 4 o 40 ng μl 1 (a e age o 16 ng μl 1 ). Fo each sample om a 384 well pla e a 5 μl aliquo was pooled and spli in o wo 2 ml Lo bind Eppendo ubes (950 μl each). To each aliquo 950 μl o AMPu e XP (Beckman, High Wycombe, UK) beads was added. Samples we e mixed, incuba ed (5 min, oom empe a u e) and placed on a magne pa icle concen a o (MPC) un il he beads we e collec ed. The supe na an was disca ded. The beads we e washed wice wi h 20 μl 70% e hanol while placed in he MPC and ai d ied (5 min). The pooled lib a y was elu ed om he beads in 17 μl o 10 mM T is HCl pH 8.0. The wo 17 μl aliquo s o he lib a y we e combined and he DNA concen a ion was de e mined using he Qbi de ice wi h he Quan -I DNA HS Assay (In i ogen). Typical DNA concen a ions we e abo e 100 ng μl 1 . The DNA size selec ion was pe o med using he Blue Pippin (Sage Science, Be e ly, MA, USA). Abou 3 μg o he lib a y in 30 μl o 10 mM T is HCl pH 8.0 and 10 μl o he R2 ladde we e sepa a ed ( igh selec ion p o ocol, 650 bp) using a 1.5% aga ose casse e acco ding o he manu ac u e ’s ins uc ions (Sage Science, Be e ly, MA, USA), he eby yielding an a e age inse size o abou 485 bp. Size selec ed samples we e collec ed in 40 μl o TRIS- TAPS bu e , pH 8.0 (Sage Science, Be e ly, MA, USA). The a e age size o he lib a y was de e mined using a High Sensi i i y Chip and an Agilen 2100 Elec opho esis Bioanalyze (Agilen ). The DNA concen a ion was measu ed using he Qbi de ice and he Quan -I DNA HS Assay (In i ogen). Size selec ed lib a ies we e quan ified using he Kappa Biosciences Illumina lib a y qPCR quan ifica ion ki (Kapa Biosciences) on a S ep One qPCR machine (The moFishe ) acco ding o he manu ac u e ’s ins uc ions and compa ed agains a known concen a ion o a PhiX con ol lib a y. Se e al lib a ies we e pooled o sequencing in an equimola manne , and he final pool was e-quan ified o sequencing ela i e o a s anda d lib a y o a known concen a ion using he Kapa Biosciences Illumina lib a y qPCR quan ifica ion ki . Sequencing- by-syn hesis o 6,144 BACs om ch omosome 2H was pe o med using an Illumina HiSeq2000 de ice (2 × 100 cycles pai ed-end, single indexing ead, 384 BACs/lane) acco ding o manu ac u e ’s ins uc ions, he eby yielding a leas 32 Gb/lane and an a e age sequence co e age o a leas 500- old pe BAC. The emaining BAC clones om 2H (384 BACs/lane) and 0H (2304 BACs/lane) we e sequenced wi h a HiSeq2500 machine (2 × 150 cycles pai ed-end, dual indexing, apid mode, yield: a leas 30 Gb/lane) using a sligh ly adap ed p o ocol wi h an addi ional no maliza ion s ep p io o sample pooling. B iefly, a cus om panel o 48 P5 and 48 P7 adap e oligos wi h 9 bp ba codes (wi h ≥4 hamming dis ance) was designed o indi idually label up o 2,304 (48 × 48) lib a ies by dual indexing. A mix u e o 2μl o BAC DNA, 0.5 μl Nex e a 10 × Reac ion Bu e , 0.1 μl Nex e a Enzyme and 2.4 μl wa e was incuba ed (5 min, 55 °C). Tn5 dena u a ion, eac ion clean-up, washing, elu ion and ans e o a esh 384 well pla e we e as desc ibed o he single-indexing lib a ies. 5 μl pu ified, agmen ed DNA, 2 μlo 5 × Kapa Robus 2G B Reac ion bu e , 0.2 μl o 10 mM dNTPs, 0.05 μl o Kapa Robus 2G Taq polyme ase, 1 μl2μM P5 p ime , 1 μl2μM P7 p ime we e combined ( eac ion olume o 10 μl) and amplified acco ding o ollowing he mal cycling p ofile: 72 °C o 3 min, 95 °C o 1 min, ollowed by www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 5 16 cycles o 95 °C o 10 s, 65 °C o 20 s and 72 °C o 3 min. The size p ofile and quan i y was de e mined as desc ibed o single-indexing lib a ies. Amplified lib a ies we e no malised using MagQuan bead echnology (GC Bio ech, Ne he lands) on a Calipe Zephy Robo (Pe kin Elme ), essen ially as desc ibed by he manu ac u e . No malised lib a ies we e elu ed in 10 μl o 10 mM T is HCl pH 8.0 and ans e ed o a esh 384 well PCR pla e.5 μl o 384 no malized samples we e pooled ( o al olume 1,920 μl). Pu ifica ion using AMPu e XP beads, washing, elu ion, size-selec ion (Blue Pippin) and quali y checks p io o sequencing we e essen ially as desc ibed o single indexing lib a ies. Sequencing-by-syn hesis o pooled lib a ies (2,304 BACs) was pe o med using an Illumina HiSeq2500 de ice ( apid un mode, 2 × 150 cycles pai ed-end, dual indexing eads) acco ding o manu ac u e ’s ins uc ions. A leas 40 Gbp/lane, and an a e age sequence co e age o >100- old pe BAC we e ob ained (Da a Ci a ion 22, Da a Ci a ion 23, Da a Ci a ion 24, Da a Ci a ion 25). Ma e-pai sequencing o MTP BACs BAC clones we e inocula ed as desc ibed o he p epa a ion o sho gun lib a ies. The bac e ial cul u es we e g own o 6 h a 37 °C in a shaking incuba o a 200 .p.m., and 384 clones we e pooled. The pool was used o inocula e 250 ml 2 × YT media supplemen ed wi h chlo amphenicol (12.5 μgml 1 ). The cul u es we e incuba ed (18 h, 37 °C, 200 .p.m.). Cells we e ha es ed by cen i uga ion (3,220 g, 20 min, 4 °C), and he supe na an was disca ded. Alkali lysis and DNA isola ion s eps we e pe o med using he La ge Cons uc ki (Qiagen, UK) essen ially ollowing he manu ac u e ’s ins uc ions. The DNA was esuspended in 4.75 ml Bu e Ex, 100 μl 100 mM ATP (Fishe Scien ific, UK) we e added and con amina ing E. coli DNA was emo ed using 150 μl ATP dependen Exonuclease (Qiagen). Du ing he incuba ion (1 h, 37 °C) a Qiagen Tip-100 column (Qiagen) was equilib a ed in Bu e QBT (Qiagen). 5 ml o Bu e QS we e added o he DNA, and he sample was applied o he equilib a ed column. The column was washed wice wi h 10 ml o Bu e QC (Qiagen). The DNA was elu ed wi h 7.5 ml o p e-wa med (65 °C) Bu e QF (Qiagen). The DNA was p ecipi a ed by adding 0.7 × olume o oom empe a u e isop opanol and cen i uga ion (20 min, 3,220 g, 4 °C). The pelle was washed wice wi h 70% e hanol, ai d ied and dissol ed in 200 μl TE bu e acco ding o manu ac u e ’s guidelines. The DNA concen a ion was measu ed using a Qubi Fluo ome e (The mo Fishe , Camb idge, UK) and adjus ed wi h wa e o 13 ng μl 1 . Fo agmen a ion 200 μl dilu ed DNA we e equilib a ed (6 min, 55 °C) and subsequen ly p o ided wi h 52 μl 5 × Tagmen Bu e Ma e-Pai and 8 μl Ma e-Pai Tagmen a ion Enzyme (Illumina, San Diego, USA). A e he incuba ion (30 min, 55 °C), 65 μl Neu alize Tagmen Bu e (Illumina, San Diego, USA) we e added, and he eac ion was incuba ed (5 min, oom empe a u e). One olume CleanPCR beads (GC Bio ech, Alphen aan den Rijn, The Ne he lands) was added, and he DNA was pu ified using magne ic sepa a ion. The DNA was elu ed in 170 μl o nuclease- ee wa e , quan ified using a Qubi fluo ome e (DNA HS assay, In i ogen) and analysed using he Agilen Bioanalyse (DNA 1,200 chip, Agilen , S ockpo , UK). S and displacemen was pe o med by combining 105.3 μl o agmen ed DNA, 13 μl 10x S and Displacemen Bu e (Illumina), 5.2 μl dNTPs (Illumina), 6.5 μl S and Displacemen Polyme ase (Illumina) and incuba ion (30 min, oom empe a u e). CleanPCR beads (0.75 olume) we e added and he DNA was pu ified using a magne . The DNA was elu ed in 30 μl nuclease- ee wa e . The concen a ion was measu ed (Qubi , DNA HS assay, In i ogen), and a 1:6 dilu ed sample was analysed using he Agilen Bioanalyse (DNA 1,200 chip, Agilen , S ockpo , UK). Size selec ion was pe o med using a Pippin Blue (Sage Science, Be e ly, MA, USA). 30 μl DNA we e p o ided wi h 10 μl loading bu e and sepa a ed on a 0.75% aga ose casse e (size selec ion cen e ed a 7 kb and collec ion be ween 6–8 kb) acco ding o he manu ac u e ’s ins uc ions (Sage Science, Be e ly, MA, USA). Size selec ed samples we e collec ed in 40 μlo TRIS- TAPS bu e (pH 8.0) (Sage Science, Be e ly, MA, USA), and analysed using he Agilen Bioanalyse (high sensi i i y chip, Agilen , S ockpo , UK) o de e mine he final lib a y size. The DNA concen a ion was measu ed using he Qubi de ice and he Quan -I DNA HS Assay (In i ogen). Ci cula isa ion was pe o med by combining 40 μl size selec ed DNA, 12.5 μl 10 × ci cula isa ion bu e (Illumina), 3 μl Ci cula isa ion Enzyme (Illumina) and 75 μl nuclease- ee wa e . The eac ion was incuba ed a 30 °C o e nigh . Linea DNA was diges ed by adding 3.75 μl Exonuclease (Illumina) and incuba ion (30 min, 37 °C). The enzyme was inac i a ed by hea (30 min, 70 °C) and he addi ion o 5 μl s op liga ion (Illumina). Ci cula ised DNA (130 μl) was shea ed in a Co a is Mic oTube AFA Fibe (P e-sli , Snap-cap, 6 × 16 mm; 2 cycles o 37 s, 10% du y cycle, 200 cycles pe bu s , 4 in ensi y, 4 °C) using he Co a is S2 de ice (Co a is, Massachuse s, USA). M280 Dynabeads (The mo Fishe ) we e p epa ed as desc ibed (Illumina). 130 μl washed M280 beads we e added o he agmen ed DNA, mixed and placed on a lab o a o (20 min, oom empe a u e). Lib a y molecules we e a fini y pu ified and washed as desc ibed (Illumina). The beads we e esuspended in a mix u e o 85 μl nuclease ee wa e , 10 μl 10x End Repai Reac ion Bu e (Ilumina) and 5 μl end epai enzyme mix (Illumina) and incuba ed (30 min, 30 °C). End epai ed lib a y molecules bound o M280 beads we e washed as desc ibed (Illumina). A-Tailing and adap e liga ion we e pe o med acco ding o manu ac u e ’s ins uc ions (Illumina). Fo PCR amplifica ion, he beads we e esuspended in a eac ion mix u e (20 μl nuclease- ee wa e , 25 μl 2x Kappa HiFi (Kappa Biosys ems, London, UK), 5 μl Illumina P ime Cock ail) and amplified (98 °C o 3 min, 12 cycles o 98 °C o 10 s, 60 °C o 30 s, 72 °C o 30 s ollowed by 72 °C o 5 min and s o age o he sample a 4 °C). Beads we e emo ed by magne ic sepa a ion and 45 μl o he p oduc s we e ans e ed o a 2 ml DNA Lobind Eppendo ube. The DNA was p ecipi a ed by addi ion www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 6 o 31.5 μl CleanPCR beads (GC Bio ech, Alphen aan den Rijn, The Ne he lands). The beads we e washed wice wi h 100 μl 70% e hanol, and he final lib a y was elu ed in 20 μl esuspension bu e (GC bio ech). The DNA concen a ion was de e mined (Qubi , DNA HS assay, In i ogen), ollowed by analysis using he Agilen Bioanalyse (High sensi i i y chip, Agilen , S ockpo , UK). Up o 12 ma e-pai lib a ies we e pooled in an equimola manne and measu ed using he Kappa qPCR Illumina quan ifica ion ki . Sequencing-by-syn hesis o pooled ma e-pai lib a ies was pe o med using an Illumina HiSeq2500 de ice ( apid un mode, 2 × 150 cycles pai ed-end, single indexing eads) acco ding o manu ac u e ’s ins uc ions (Da a Ci a ion 26, Da a Ci a ion 27). Sequence assembly o indi idual BACs Assembly o gene-con aining BACs (UCR/JGI). A o al o 15,661 gene-bea ing BACs we e pai ed-end sequenced (2 × 100 cycles) using he Illumina HiSeq2000 pla o m (Illumina, Inc., San Diego, CA, USA) applying a combina o ial pooling design 24 , as desc ibed in Munoz-Ama iain e al. 13 . Reads we e quali y immed, decon olu ed, and hen assembled BAC-by-BAC using Vel e e sion 1.2.09 ( e . 25) wi h he pa ame e k se o 45. Sequences o an addi ional 50 andomly chosen BACs included in Munoz-Ama iain e al. 13 we e de i ed using he Sange me hod by Jane G imwood (US Depa men o Ene gy Join Genome Ins i u e) and Je emy Schmu z (HudsonAlpha Ins i u e o Bio echnology), including sha e and ansposon sequencing. The assignmen o BACs o ch omosome a ms/pe i- cen ome ic egions was pe o med using CLARK 26 , an accu a e k-me -based classifica ion me hod ha is much as e han BLASTN o MegaBLAST. CLARK makes assignmen s by using a p ebuil da abase o k-me s ha a e specific o each ch omosome a m/pe i-cen ome ic egion. Assembly o MTP BACs om ba ley ch omosomes 1H, 3H, 4H, 6H and 7H (FLI and IPK). A o al o 10,148 BACs mainly o igina ing om ba ley ch omosome 3H we e sequenced on he Roche 454 sys em. Reads we e decon olu ed and assigned o indi idual BACs 16 . Reads we e quali y immed acco ding o he manu ac u e ’s ecommenda ions. Reads we e sc eened o E. coli and ec o sequences wi h MegaBLAST 27 . Assemblies we e hen cons uc ed om he clean eads using he MIRA so wa e 28 as desc ibed in S eue nagel, e al. 16 and Taudien, e al. 29 . A o al o 41,004 BACs we e sequenced on Illumina machines (mainly HiSeq2000) in pools o up o 672 indi idually ba coded BAC clones. Pai ed-end eads we e quali y immed wi h he CLC oolki and sc eened o E. coli and ec o sequences wi h MegaBLAST. Assemblies we e ob ained by unning CLC Assembly Cell Ve sion 4.0.6 be a wi h de aul pa ame e s. Con igs de i ed wi h low ead co e age as well as con igs smalle han 500 bp we e emo ed using he c i e ia desc ibed in Beie , e al. 17 . The esul an con igs we e hen compa ed o NCBI’s nucleo ide da abase using MegaBLAST o check o possible con amina ion. Con igs wi h non-plan hi s we e ei he comple ely emo ed o immed. Sca olding o MTP BACs om ba ley ch omosomes 1H, 3H, 4H, 6Hand7H(FLIandIPK).Sca olding was pe o med as desc ibed in Beie e al. 17 B iefly, ma e-pai eads we e mapped agains he conca ena ed assemblies o up o 384 BACs using BWA mem e sion 0.7.4 ( e . 30) wi h de aul pa ame e s. Only ead pai s mapping uniquely (minimal mapping quali y o Q40) o di e en con igs o he same BAC assembly we e e ained. These eads we e used o sca old indi idual BACs using SSPACE e sion 3.0 S anda d 31 . I mul iple ma e-pai lib a ies we e p esen (MiSeq ma e-pai eads as well as HiSeq2000 ma e-pai eads) an i e a i e sca olding p ocedu e 17 was used. Assembly o MTP BACs om ba ley ch omosome 5H (BGI). Ob ained aw sequence eads om 5H MTP BACs we e fil e ed o gene a e high-quali y eads by he ollowing c i e ia: (1) eads con aining mo e han 2% o Ns o wi h poly-A s uc u es we e emo ed; (2) eads wi h ≥40% low quali y bases o sho inse size lib a ies (60% o la ge inse size lib a ies) we e excluded; (3) eads con aining adap e s we e emo ed; (4) PCR duplica es we e de ec ed and excluded; (5) emo al o eads con amina ed by E. coli, ec o sequences o phage sequences. High-quali y eads we e hen used o assembly. BACs we e assembled using SOAPdeno o e sion 2.01 ( e . 32) mul iple imes using di e en k and m alues (main pa ame e in SOAPdeno o assembly). In o al each BAC was assembled 45 imes (k om 33 o 66, only odd numbe s and m om 1 o 3). The N50 was examined o each assembly and he assembly wi h he la ges N50 was e ained as he final assembly esul o each BAC. Sca olding o MTP BACs om ba ley ch omosomes 5H (BGI). Assemblies om pai ed-end sequences we e used as e e ence o mapping 2, 5 and 10 kb ma e-pai eads ob ained om ba ley genomic WGS da a wi h SOAPaligne /soap2 e sion 2.21 wi h pa ame e s –p6– 3–R. Ma e-pai ead pai s mapped in his ashion we e used in conjunc ion wi h he co esponding pai ed-end ead pai s o e-assemble each BAC using SOAPdeno o e sion 2.01 as desc ibed abo e. Assembly o MTP BACs om ba ley ch omosomes 2H and ‘0H’(EI). Minimal iling pa h BACs om (i) ba ley ch omosomes 2H o om (ii) finge p in ed con igs no assigned o ch omosomes ( e med ‘0H’) we e sequenced. A e demul iplexing, sample quali y con ol (QC) in o ma ion was gene a ed using Fas QC 33 . Con amina ion sc eening was ca ied ou using Kon aminan 34 . Reads we e sc eened using a k-me size o 21 agains a ange o po en ial con aminan s (Phi X, E. coli,En e obac e cloacae www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 7 genomic DNA and BAC ec o ) and con amina ed eads o eads wi h quali y alues o30 we e emo ed. ABySS assemble ( 1.5.1) 35 was used o assemble he fil e ed pai ed-end eads o each BAC indi idually (k-71, l-91 b-0). Pai ed-end con igs we e compa ed o NCBI’s NR da abase using BLAST o check o hi s o non-plan o ganisms using e- alue 1e-4 as h eshold. The ob ained hi s we e compa ed o NCBI axonomy using ‘ as acmd’ o ob ain common names used o check o any non-plan hi s. Sca olding o MTP BACs om ba ley ch omosomes 2H and ‘0H’(EI). Illumina Nex e a ma e-pai lib a ies we e c ea ed om pools o 384 BACs. A e quali y checking he eads using PAP 34 , he eads we e me ged using FLASH ( e sion 1.2.9) 36 . Nex clip ( 0.8) 37 was un on he flashed eads o im he junc ion adap e s. A k-me -based app oach was used o assign ma e-pai eads o indi idual BACs wi h KAT ( 1.0.4) (h ps://gi hub.com/TGAC/KAT). Sca olding and gap closing we e pe o med on each BAC indi idually using an in-house shell sc ip (a ailable om Gi Hub: h ps://gi hub.com/DhSaTGAC/ BAC-assembly-pipeline.gi ). SOAPdeno o sca olde e sion 2.01 ( e . 38) was applied o sca old he ABySS pai ed-end con igs using he k-me classified ma e-pai eads wi h pa ame e s k=41, -G 30, -F, -w and -L 100. The esul ing sca olds we e hen edi ed o eplace long s e ches (>20) o C/G wi h ‘N’cha ac e s as SOAP is known o subs i u e ‘N’s wi hin pai ed-end con igs o C/G. The sca olds we e hen passed h ough GapClose ( 1.12- 6), a SOAP2 module, o fill in long s e ches o ‘N’s p oduced du ing he sca olding s eps. Con igs and sca olds sho e han 500 bp we e emo ed o p oduce he final assembly pe BAC. Splash con amina ion checks o MTP BACs om ba ley ch omosomes 2H and ‘0H’(EI). The aw eads wi hin each pla e we e aligned o one side o he ec o sequence adjacen o he es ic ion enzyme cu si e using exone a e 39 . Subs ings o size 20 bp we e ex ac ed om aligning eads con aining he BAC sequence adjacen o he ec o sequence. Flanking sequences om each BAC we e clus e ed based on a Hamming dis anceo3 and consensus sequences gene a ed o accoun o sequencing e o s. These we e compa ed wi h neighbo ing wells o check o po en ial con amina ion caused by splash du ing lab p ocessing s eps. Whe e con amina ion be ween neighbo ing wells was indica ed, he assembled con igs om each BAC in ques ion we e aligned in a pai wise ashion using exone a e and he o al pe cen age o simila sequence (≥99% iden i y) was compu ed. In cases whe e neighbo ing BACs sha ed mo e han 10% simila sequence, bo h BACs we e esequenced. Pseudomolecule cons uc ion Ini ial con amina ion emo al. Sequence assemblies o 66,586 MTP clones, 5,468 non-MTP BACs and 15,044 gene-bea ing clones 13 ( o al numbe o unique BACs: 87,098) we e combined in o a single FASTA file (Da a Ci a ion 28,Da a Ci a ion 29,Da a Ci a ion 30). I a clone had wo o mo e independen sequence assemblies, we selec ed he one wi h he la ges N50 alue o u he analyses. BAC assemblies we e aligned o a cus om lib a y o po en ial con aminan s (Da a Ci a ion 31) including phages, bac e ial and ec o sequences using megablas 27 . Regions aligning o con aminan s (c i e ia: (alignmen leng h ≥500 bp AND iden i y ≥80%) OR (iden i y ≥90%)) we e emo ed om he assembly using UNIX sc ip s and BEDTools 40 . Sequences sho e han 500 bp o consis ing o less han 500 p ope nucleo ides (ACGT cha ac e s) a e con amina ion emo al we e disca ded. This s ep emo ed 55.5 Mb (0.5%) o he assembled BAC sequence. Sequence alignmen o BACs sequences and o e lap de ec ion. A e con amina ion emo al, a se o 87,075 BAC assemblies (Table 1, Da a Ci a ion 32) was aligned agains i sel using megablas 27 wi h a wo d size o 44, e aining only alignmen s wi h iden i y ≥99% and alignmen leng h ≥500 bp. Two se s o o e laps (s ingen and pe missi e) be ween BACs we e defined om he BLAST esul s o all BACs agains each o he . Pai s o BACs we e conside ed as po en ially o e lapping unde s ingen c i e ia i he e was a leas one high-sco ing pai (HSP) wi h alignmen leng h ≥5 kb and iden i y ≥99.8%. Unde pe missi e c i e ia, we equi ed a leas one HSP wi h alignmen leng h ≥2 kb and iden i y ≥99.5%. Fo all pai s o po en ially o e lapping BACs (unde ei he se o c i e ia), he size o hei o e lapping egions was de e mined using UNIX sc ip s and BEDTools 40 as he ex en o non- edundan egions in he BAC sequences (i.e., con igs o sca olds) con ained in HSPs ≥500 bp and iden i y ≥99.5% be ween BAC sequences ha ing a leas one HSP wi h alignmen leng h ≥5 kb and iden i y ≥99.8% (s ingen c i e ia) o alignmen leng h ≥2 kb and iden i y ≥99.5% (pe missi e c i e ia). HSPs less han 200 bp apa we e combined in o one wi h BEDTools (command ‘me ge’). BAC o e lap in o ma ion was impo ed in o he R s a is ical en i onmen 41 o use in gene ic ancho ing and me ging sequence assemblies o adjacen BAC clones (see sec ion ‘Cons uc ion o he BAC o e lap g aph’). Alignmen o BACs o he BioNano map o ba ley c . Mo ex. An op ical map o he genome o ba ley c . Mo ex was gene a ed using he I ys pla o m o BioNano Genomics using N .BspQI as he nicking enzyme. Fu he de ails o he op ical map p ocedu e a e desc ibed in Masche e al. 42 An in silico BspQI diges was pe o med wi h he Knicke s so wa e (h p://www.bionanogenomics.com) using de aul pa ame e s. Res ic ion maps o BAC sequences we e aligned o he BioNano map o ba ley c . Mo ex 42 (Da a Ci a ion 33) wi h I ysView so wa e 43 (h p://www.bionanogenomics.com) using he www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 8 command line ool Re Aligne ( e sion 3827) wi h he ollowing pa ame e s ‘-M 2 -T 1e-4 -ex end 1 -biasw 0’ o epo all alignmen s wi h a confidence sco e ≥4. Cons uc ion o he upda ed POPSEQ map o he Mo ex x Ba ke mapping popula ion. An ul a- dense linkage map had been cons uc ed p e iously 5 by shallow whole-genome sho gun sequencing o 90 ecombinan inb ed lines (RILs) de i ed om a c oss be ween he ba ley cul i a s Mo ex and Ba ke. We wished o inc ease he esolu ion o his map by educing he a e age ac ion o missing da a pe SNP ma ke . Towa ds his aim, we sequenced he exis ing Illumina pai ed-end lib a ies o 87 RILs o highe co e age (2–3x) and combined hem (Da a Ci a ion 34) wi h he exis ing ead da a se 5 (ENA accession: ERP002184). Map cons uc ion ollowed he p ocedu es desc ibed in Chapman e al. 44 . Reads we e aligned o he whole-genome sho gun assembly o ba ley c . Mo ex 4 (NCBI accession: CAJW01) wi h BWA mem e sion 0.7.5a ( e . 45). So ing, con e sion o BAM o ma and emo al o duplica e eads was done wi h Pica dTools e sion 1.100 (h p://b oadins i u e.gi hub.io/pica d/). Va ian de ec ion and geno ype calling we e pe o med wi h SAMTools e sion 0.1.19 (commands ‘sam ools mpileup –BD’and ‘bc ools iew –c g’). The esul an VCF file was fil e ed using an AWK sc ip (Supplemen a y Tex S3 o Masche e al. 2013 ( e . 46)). Homozygous geno ype calls we e se o missing i hei ead dep h was 0 o hei geno ype quali y below 3. He e ozygous geno ype calls we e se o missing i hei ead dep h was below 3 o hei geno ype quali y below 5. Va ian s wi h (i) a quali y sco es below 40, (ii) mo e han 10% he e ozygous geno ype calls, (iii) mo e han 90% missing da a a e geno ype call fil e ing, o (i ) a mino allele equency below 5% we e disca ded. SNP in o ma ion was agg ega ed a he con ig le el o de i e consensus geno ypes as desc ibed in he sec ion ‘F amewo k map cons uc ion’in he Me hods sec ion o Chapman e al. 44 Fo map cons uc ion wi h MSTMap 47 , he popula ion ype ‘RIL8’was used. Addi ional con igs we e inse ed in o he amewo k map as desc ibed in Chapman e al. 44 (sec ion ‘Ancho ing sca olds on o he amewo k map’) using p e iously published ead da a 5 . Va ian calling and map cons uc ion we e done o he O egon Wol e Ba ley (OWB) doubled haploid popula ion using he same p ocedu es wi h he ollowing wo changes: (i) he e ozygous geno ype calls we e excluded and (ii) he popula ion ype ‘DH’was used o map cons uc ion wi h MSTMap 47 . Map posi ions in he OWB map we e in e pola ed in o he Mo ex x Ba ke map using loess eg ession in R 41 . A consensus posi ion was de i ed as ollows: i map posi ions disag eed by mo e han 5 cM in bo h maps, a con ig was conside ed unancho ed; o he wise, he Mo ex x Ba ke posi ion was p e e ed i a ailable. The final map assigned gene ic posi ions o 791,176 WGS con igs (Table 2, Da a Ci a ion 35), compa ed o 723,499 ancho ed con igs in he o iginal POPSEQ map 5 . Gene ic ancho ing o single BAC clones. The gene ic posi ions o Mo ex WGS con igs in he upda ed POPSEQ map we e li ed o BAC sequences ia sequence alignmen . The se o all con igs o he whole- genome sho gun assembly o ba ley c . Mo ex 4 (NCBI accession: CAJW01) was aligned o all BAC assemblies wi h megablas 27 using a wo d size o 44 and e aining only alignmen s wi h iden i y ≥99.8% and alignmen leng h ≥1,000 bp. Fo each BAC clone, he gene ic posi ions o WGS con igs aligning o i s cons i uen sequences we e abula ed and a gene ic posi ion o a clone was de i ed using a majo i y ule wi h unc ions o he R package ‘da a. able’(h ps://c an. -p ojec .o g/web/packages/da a. able/index. h ml). Nine y pe cen o con igs assigned o a BAC had o o igina e o he majo ch omosome and he s anda d de ia ion o gene ic posi ions had o be ≤3 cM. BACs wi hou alignmen s o ancho ed WGS con igs we e conside ed as unancho ed; hose no mee ing he consis ency c i e ia we e flagged as ‘inconsis en ly ancho ed’. In he second s ep, unancho ed clones we e posi ioned by u ilizing posi ional in o ma ion om neighbo ing BACs. We conside ed as neighbo s o a gi en clone B all hose BACs ha o e lapped o a leas 10% o hei assembled leng hs wi h clone B. The gene ic posi ion o an MTP ch omosome no o . BACs in MTP no. o sequenced BACs no. o ancho ed BACs*a e age no. o sequences a e age N50 (kb) 1H 6,993 6,983 (99.9%) 6,410 (91.8%) 7.6 81.2 2H 9,061 8,969 (99.0%) 8,195 (91.4%) 9.9 104.5 3H 8,841 8,807 (99.6%) 8,303 (94.3%) 7.7 87.5 4H 8,314 8,306 (99.9%) 7,783 (93.7%) 6.7 91.2 5H 8,426 8,358 (99.2%) 7,573 (90.6%) 9.7 72.2 6H 8,305 7,886 (95.0%) 6,476 (82.1%) 7.4 70.7 7H 8,576 7,970 (92.9%) 6,842 (85.8%) 8.5 65.5 ‘0H’ † 8,256 8,031 (97.3%) 6,714 (83.6%) 7.6 83.6 Non-MTP —21,765 20,397 (93.7%) 14.5 33.7 To al 66,772 87,075 78,693 (90.4%) 9.8 70.3 Table 1. BAC assembly and ancho ing s a is ics. *Numbe and pe cen age o BAC clones ha ha e been assigned gene ic posi ions in he POPSEQ map. † BAC clones in physical con igs ha had no been assigned o ch omosomes. www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 9 leng h ≥500 bp and an alignmen iden i y ≥99.5%. Regions o he old non- edundan sequence co e ed by C (as de e mined by commands o BEDTools 40 sui e) we e emo ed and con ig C was added ins ead. This p ocedu e was pe o med o each C2- ype con ig. Nex , we que ied he GMAP alignmen s o genes ha had no alignmen s o he non- edundan sequence, bu we e ep esen ed ei he in (a) he Mo ex WGS con igs o in (b) sequences o BACs excluded om he o e lap analysis. We conside ed sequence o ype (a) and (b) as ‘addi ional gene-bea ing sequences’. We aligned hese addi ional gene-bea ing sequences o he non- edundan sequence wi h megablas 27 using a wo d size o 44 and conside ing only high-sco ing pai s wi h an alignmen leng h ≥500 bp and an alignmen iden i y ≥99.5%. Regions co e ed by he non- edundan sequence unde hese alignmen c i e ia we e sub ac ed om he addi ional gene-bea ing sequences and sequence agmen s wi h a leng h ≥500 bp we e added o he non- edundan sequence. Final con amina ion emo al. We iden ified egions in he non- edundan sequence ha we e no co e ed by whole-genome sho gun eads o c . Mo ex. Alignmen o WGS eads and ead dep h calcula ion we e done as desc ibed in he sec ion ‘Alignmen o Hi-C da a o es ic ion agmen s’. Regions o he non- edundan sequence no co e ed by Mo ex WGS eads and wi h a leng h ≥500 bp we e ex ac ed using UNIX command line ools and BEDTools 40 (command ‘ge as a’). The ex ac ed sequences we e aligned o he NCBI NT da abase wi h megablas 27 using a wo d size o 44 and equi ing he high-sco ing pai s o ha e a leng h o a leas 100 bp and an alignmen iden i y ≥80%. We e ained only hi s whose desc ip ion in he NCBI NT da abase did no ma ch he ollowing egula exp ession (R syn ax) ep esen ing a lis o common and axonomic names o plan species: ‘Ho deum|T i i|Populus|Aegilops|A ena|Alnus|A .squa osa|Mo us|Nelumbo|B assica|Cucumis|Ci us| Camelina|F aga ia|Lo us|Ta enaya|Spa ina|Eucommia|So ghum|Co ylus|Theob oma|Phaseolus|Ba ley| T i olium|Elymus|B achypodium|Be a ulga is|Ricinus|Licania|Phoenix|H . ulga e|Py us|Malus|P unus| Saccha um|Hype icum|Whea |O yza|hlo oplas |Secale|Vi is|Que cus’ Regions o e lapping he BLAST hi s passing hese fil e s we e cu om he non- edundan sequence wi h BEDTools 40 (command ‘sub ac ’). Sequences sho e han 500 bp a e he emo al o con aminan sequences we e disca ded. This s ep emo ed 5 Mb (0.1%) o he assembled sequence. Cons uc ion o pseudomolecule sequences o ch omosome 1H—7H and ch Un. We cons uc ed pseudomolecules o he se en ba ley ch omosomes by placing he sequence agmen s o single BAC assemblies ha cons i u e he non- edundan sequence acco ding o he Hi-C map posi ions o he BAC o e lap clus e s hese agmen s belong o. Sequences no ancho ed by Hi-C we e placed on ch Un (‘ch omosome unassigned’). The o de o clus e s was aken om he Hi-C map. BACs wi hin he same clus e we e o de ed acco ding o he minimum spanning ee o he BAC o e lap g aph o he clus e and o ien ed ela i e o he elome es using he Hi-C o ien a ion o he clus e i a ailable. The ela i e o de o sequence agmen s o igina ing om he same BAC bin (see sec ion ‘Cons uc ion o he BAC o e lap g aph’) could no be de e mined so ha he placemen o sequences wi hin a BAC bin (a e age size: 70 kb) is a bi a y. Ch Un is composed o (i) sequence agmen s o igina ing om BAC o e lap clus e s no placed in he Hi-C map, o (ii) gene-bea ing agmen s o BAC sequences and Mo ex WGS con igs selec ed in addi ion o he non- edundan sequence (see sec ion Iden ifica ion o addi ional gene- bea ing sequences). A gap o 100 N cha ac e s was inse ed be ween adjacen sequence agmen s. Pseudomolecules o all ch omosomes and ch Un we e combined in o a single FASTA file (Da a Ci a ion 42). To accommoda e limi a ions o he Sequence/Alignmen Map o ma (see Usage No es) spli pseudomolecules wi h a size below 512 Mb we e cons uc ed by b eaking pseudomolecules a bi a ily a b eaks be ween sequence con igs (Da a Ci a ion 43, Da a Ci a ion 44). A BED file indica ing he placemen o BAC sequence agmen s, Mo ex WGS con igs and in e cala ing gaps in he (spli ) pseudomolecules is a ailable o download (Da a Ci a ion 45, Da a Ci a ion 46). A abula summa y o he posi ional in o ma ion inco po a ed in o pseudomolecules is gi en in Da a Ci a ion 41. Masking o esidual edundancy Residual edundancy a ising om unde ec ed o e laps be ween adjacen BACs was de ec ed and masked by aligning he pseudomolecules sequence o i sel wi h megablas 27 . Genomic in e als con ained in BLAST hi s wi h a leng h ≥5 kb and an iden i y ≥99.8% we e conside ed as po en ially edundan (PR) egions. PR egions we e classified o decide which sequence o a edundan pai o mask: (i) PR egions assigned o ch omosomal pseudomolecules (as opposed o ch Un), bu ha ing BLAST hi s only o o he ch omosomes we e conside ed as o igina ing om chime ic BAC assemblies inco po a ing un ela ed sequences om di e en ch omosomes and masked wi h Ns; (ii) an analogous p ocedu es was used o find in ach omosomal chime as based on Hi-C map in o ma ion; (iii) PR egions on ch Un ha had alignmen s o egions on ch omosomal pseudomolecules we e masked, (i ) o o he PR egions one sequence o a edundan pai was chosen a bi a ily. Posi ions o masked egions on he (spli ) pseudomolecules we e w i en in o a BED file (Da a Ci a ion 47, Da a Ci a ion 48). Masking was done wi h BEDTools 40 (command ‘mask’) o e w i ing nucleo ides in edundan in e als wi h N cha ac e s. Masked e sions o he (spli ) pseudomolecules a e p o ided as Da a Ci a ion 49, Da a Ci a ion 50). www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 16 POPSEQ gene ic map based on pseudomolecule sequence A e he cons uc ion o he map-based e e ence sequence, we cons uc ed an upda ed high- esolu ion gene ic map o he Mo ex x Ba ke popula ion o alida e he o de o gene ic map in he e e ence Figu e 2. Collinea i y be ween he Hi-C map and wo gene ic maps. The posi ions o gene ic ma ke s (x-axis) a e plo ed agains hei gene ic posi ions (y-axis) in a GBS map ( op ow) and a POPSEQ map (bo om ow) o he Mo ex x Ba ke ecombinan inb ed lines. www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 17 sequence. Raw eads (see sec ion ‘Cons uc ion o he upda ed POPSEQ map o he Mo ex x Ba ke mapping popula ion’) we e aligned o he ba ley pseudomolecules wi h BWA mem ( e sion 0.7.12) 45 . Checking ma ed mapped pai ed eads, so ing, con e sion o BAM o ma and ma king o duplica e ead pai s we e done wi h Pica dTools e sion 2.300 (h p://b oadins i u e.gi hub.io/pica d/). Va ian de ec ion and geno ype calling we e pe o med using GATK Toolki e sion 3.3.0 (command ‘Haplo ypeCalle ’) 57 . A o al o fi e RILs wi h >3% he e ozygous a ian s we e emo ed. A a ian posi ion was emo ed i mo e han 10% o all samples we e called he e ozygous, he e we e mo e han 80% missing da a, o he mino allele equency (in he non-missing da a) was smalle han 5%. SNP in o ma ion was agg ega ed a he con ig le el o de i e consensus geno ype blocks wi h alse disco e y a e calcula ed based on he quali y o each a ian call in he block. High-confidence geno ype blocks we e ob ained based on a Bon e oni co ec ion h eshold. Gi en he ac ha he leng h o c osso e ac s is significan ly la ge han ha o non-c osso e ac s and non-c osso e ac s would enla ge he gene ic dis ance a ificially, we only e ained high-confidence geno ype blocks wi h mo e han 1 Mb ac leng h, which a e likely o be de i ed om c osso e s. Rep esen a i e non- edundan genomic a ian s o high-confidence geno ype blocks we e ex ac ed and used o he cons uc ion o a high- esolu ion map h ough MSTMap 47 . We u he ancho ed all emaining ma ke s o he gene ic map by he C p og am ‘cancho ’ 5 . The final POPSEQ map consis ed o 9,012,742 SNP a ian s defined on he pseudomolecule sequence Da a ci a ion 51). Rep esen a ion o ull-leng h cDNAs The ep esen a ion o gene models in he whole-genome genome assembly o ba ley c . Mo ex 4 and in he pseudomolecules was compa ed by aligning a se o 22,651 publicly a ailable ull-leng h cDNAs 55 o he assemblies using he GMAP splice aligne so wa e 56 . The GMAP alignmen ou pu was hen fil e ed. I a ull-leng h cDNA had mul iple hi s, only he hi wi h he highes % iden i y was conside ed. Hi s we e u he fil e ed by iden i y (≥98%) and co e age ( ≥95%). This esul ed in a se o hi s ep esen ing genes eco e ed in ac on a single genomic con ig/ch omosome. Code a ailabili y R and shell sou ce code o he cons uc ion o he BAC o e lap g aph and he Hi-C map is p o ided as Da a Ci a ion 52. Code can be e-used unde he e ms o he MIT license. Da a Reco ds BAC sequence aw da a was submi ed o he Eu opean Nucleo ide A chi e (ENA) (Da a Ci a ion 1, Da a Ci a ion 2, Da a Ci a ion 3, Da a Ci a ion 4, Da a Ci a ion 5, Da a Ci a ion 6, Da a Ci a ion 7, Da a Ci a ion 8, Da a Ci a ion 9, Da a Ci a ion 10, Da a Ci a ion 11, Da a Ci a ion 12, Da a Ci a ion 13, Da a Ci a ion 14, Da a Ci a ion 15, Da a Ci a ion 16, Da a Ci a ion 17, Da a Ci a ion 18, Da a Ci a ion 19, Da a Ci a ion 20, Da a Ci a ion 21, Da a Ci a ion 22, Da a Ci a ion 23, Da a Ci a ion 24, Da a Ci a ion 25, Da a Ci a ion 26, Da a Ci a ion 27). BAC assemblies we e submi ed o ENA o NCBI (Da a Ci a ion 28, Da a Ci a ion 29). Raw da a o POPSEQ (Da a Ci a ion 35), GBS (Da a Ci a ion 38) and Hi-C mapping (Da a Ci a ion 40) we e submi ed o ENA. P ocessed da ase s a e accessible as Figu e 3. Collinea i y be ween he Hi-C map and a cy ogene ic map o ch omosome 3H. Do s ma k he posi ions o p obes in he cy ogene ic map (x-axis) and he Hi-C-de i ed pseudomolecule (y-axis). A linea eg ession line ( ed) was fi ed wi h he R unc ion lm(). No e ha cy ogene ic da a is no a ailable o dis al egions because p obes we e designed only o non- ecombining pe i-cen ome ic egions 61 . www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 18 QRLPDAAGSSAEEHSGQDKLLIVVTPTR ARASQAYYLSRMGQTLRLVRPPVLWVVV EAGKPTPEAALELRRTAVMHRYVGCCDA LNASASPAVDFRPHQLNAGLEVVENHRL DGVVYFADEEGVYSLPLFDRLRQIRRFG TWPVPTISDGGHGVVLEGPVCKQNQVVG WHTSGDANKLQRFHVAMSGFAFNSTMLW DPRLRSHKAWNSIRHPEMVEQGFQGTTF VEQLVEDESQMEGIPADCSQIMNWHVPF GSESPVYPKGWRSAANLDVIIPLK Figu e 4. Accessing sequence and posi ional in o ma ion wi h he ba ley genome explo e (BARLEX). The ba ley pseudomolecule da a was impo ed in o BARLEX, whe e i is di ec ly linked o he IPK Ba ley BLAST se e . Use s can pas e a nucleo ide o amino acid sequence (1) in o he BARLEX inpu que y o m and selec e e ence da abase such as pseudomolecules sequence, he se o all BAC assemblies o anno a ed genes (2). The sequence is hen ans e ed o he IPK ba ley BLAST Se e (3). The web page wi h he BLAST esul s (4) con ains e e ences o BARLEX in o ma ion pages o di e en s uc u al uni s (BAC sequence con igs, BAC, BAC clus e , ch omosomal Hi-C map). Fo example, he pages o BAC sequence con igs isualize he epea con en based on genome-wide k-me his og ams (5) and a e linked o a g aph-based isualiza ion (6) o he en i e BAC assembly. Summa y s a is ics and posi ional in o ma ion o BAC clus e s a e p esen ed in ables ha can be sea ched, so ed and subse ed using use -defined c i e ia (7). Use s can con e pseudomolecule coo dina es (AGP posi ions) o in e als in he unde lying BAC sequence assemblies (8). www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 19 Digi al Objec Iden ifie s (DOIs) in he Plan Genomics and Phenomics Resea ch Da a Reposi o y 58 (Da a Ci a ion 30, Da a Ci a ion 31, Da a Ci a ion 32, Da a Ci a ion 33, Da a Ci a ion 34, Da a Ci a ion 36, Da a Ci a ion 37, Da a Ci a ion 39, Da a Ci a ion 41, Da a Ci a ion 42, Da a Ci a ion 43, Da a Ci a ion 44, Da a Ci a ion 45, Da a Ci a ion 46, Da a Ci a ion 47, Da a Ci a ion 48, Da a Ci a ion 49, Da a Ci a ion 50, Da a ci a ion 51, Da a Ci a ion 52). DOIs we e egis e ed wi h e!DAL 59 . Technical Valida ion Collinea i y be ween gene ic maps and pseudomolecules To alida e he o de o sca olds in he Hi-C map, we compa ed he o de o gene ic ma ke loci in he Hi-C-de i ed pseudomolecules o hei posi ions in linkage maps. Fi s , we used geno yping-by- sequencing (GBS) 11,50 o ype single-nucleo ide polymo phisms (SNPs) seg ega ing in a bi-pa en al popula ion comp ising 2,398 ecombinan inb ed lines (RILs). A o al o 2,637 SNPs we e de ec ed by aligning GBS eads and calling a ian s and geno ypes using a p e iously published pipeline 46 . Second, we eanalysed WGS e-sequencing da a o a subse o he same popula ion (POPSEQ da a) comp ising 90 RILs. Cons uc ion o a amewo k linkage map and inse ion o addi ional ma ke s we e pe o med essen ially as desc ibed by Chapman e al. 44 A do plo compa ison o physical and gene ic SNP posi ions e ealed ha ma ke o de s we e highly collinea be ween he pseudomolecules and bo h he GBS and POPSEQ map o he Mo ex x Ba ke popula ion (Fig. 2). Collinea i y be ween a cy ogene ic map and he pseudomolecule o ch omosome 3H We could no alida e he o de o BAC o e lap clus e s in he la ge pe i-cen ome ic egions because o se e ely ep essed ecombina ion 3,60 . The e o e, we compa ed he o de o p obes mapped by fluo escence in-si u hyb idiza ion o ch omosomal loca ions on ch omosome 3H and hei co esponding sequences in he pseudomolecule o 3H. Since p obes we e de i ed om BAC sequences associa ed wi h physical con igs, hei posi ion om he e e ence sequence could be de e mined om he BAC o e lap g aph. The compa ison showed ha he cy ogene ic and Hi-C maps we e highly collinea in pe i- cen ome ic egions o ch omosome 3H (Fig. 3). Rep esen a ion o ull-leng h cDNAs To assess he comple eness o ou assembly, we checked o he p esence o high-confidence ansc ip sequences. The ep esen a ion o gene models in he whole-genome sho gun assembly o ba ley c . Mo ex 4 and in he map-based e e ence assembly was compa ed by aligning a se o 22,651 publicly a ailable ull-leng h cDNAs 55 o ba ley c . ‘Ha una Nijo’. A e aligning and fil e ing, 18,062 (79.74%) in ac ull-leng h cDNAs we e ound in he pseudomolecules, whe eas only 10,496 (46.33%) we e eco e ed in he whole-genome assembly. This inc ease in he numbe o co ec ly ep esen ed ull-leng h cDNAs indica es he e o in es ed in he map-based assembly. Ne e heless, a significan p opo ion o genes emain agmen ed e en in he pseudomolecule assembly (20.26%), and p esumably hese la gely ep esen di ficul o assemble genes ha con ain e.g., mic osa elli es, long homopolyme s e ches and o he di ficul ea u es, and/o o m pa o complex gene amilies ha a e di ficul o esol e. I is likely ha only longe ead echnologies such as Pacific Biosciences (h p://www.pacb.com) o Ox o d Nanopo e (h ps://www.nanopo e ech.com) will be able o esol e hese mo e di ficul cases. Fu he esul s on gene space comple eness based on an au oma ed gene anno a ion o he pseudomolecules, and on he ep esen a ion o epe i i e elemen s a e desc ibed elsewhe e 42 . Usage No es Posi ional in o ma ion o BAC sequences, physical con igs and WGS con igs can be accessed ia he ba ley genome explo e BARLEX (Fig. 4). BLAST sea ches agains he ba ley pseudomolecules can also be ca ied ou in BARLEX. We no e ha p ocessing BAM files wi h sho ead alignmen s o he ull pseudomolecules wi h commonly used ools such as SAM ools 52 o BEDTools 40 may no wo k as expec ed because o es ic ions on he ch omosome size (512 Mb) o indexing file in Sequence Alignmen /Map (SAM) o ma 52 . To ci cum en his issue, we ha e spli he pseudomolecules in o wo pa and p o ide (i) a FASTA file wi h spli pseudomolecules (Da a Ci a ion 44) along he wi h he in ac sequences and (ii) a BEDfile o con e be ween ull and spli pseudomolecule coo dina e (Da a Ci a ion 43) Al e na i ely, he CRAM o ma (h ps://sam ools.gi hub.io/h s-specs/CRAM 3.pd ) may be used ins ead o he BAM o ma . We no e ha he o ien a ion o sequence con igs wi hin indi idual BACs in he pseudomolecules is a bi a y, hus he o de and o ien a ion o sequences in he pseudomolecules is accu a e only up o esolu ion o ~100 kb. Re e ences 1. Schul e, D. e al. The in e na ional ba ley sequencing conso ium--a he h eshold o e ficien access o he ba ley genome. Plan physiology 149, 142–147 (2009). 2. Schul e, D. e al. BAC lib a y esou ces o map-based cloning and physical map cons uc ion in ba ley (Ho deum ulga e L). BMC genomics 12, 247 (2011). 3. A iyadasa, R. e al. A sequence- eady physical map o ba ley ancho ed gene ically by wo million single-nucleo ide poly- mo phisms. Plan physiology 164, 412–423 (2014). 4. In e na ional Ba ley Genome Sequencing Conso ium. A physical, gene ic and unc ional sequence assembly o he ba ley genome. Na u e 491, 711–716 (2012). www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 20 5. Masche , M. e al. Ancho ing and o de ing NGS con ig assemblies by popula ion sequencing (POPSEQ). The Plan Jou nal 76, 718–727 (2013). 6. Lande , E. S. e al. Ini ial sequencing and analysis o he human genome. Na u e 409, 860–921 (2001). 7. Schnable, P. S. e al. The B73 maize genome: complexi y, di e si y, and dynamics. Science 326, 1112–1115 (2009). 8. Lam, E. T. e al. Genome mapping on nanochannel a ays o s uc u al a ia ion analysis and sequence assembly. Na u e bio echnology 30, 771–776 (2012). 9. Liebe man-Aiden, E. e al. Comp ehensi e mapping o long- ange in e ac ions e eals olding p inciples o he human genome. Science 326, 289–293 (2009). 10. Bu on, J. N. e al. Ch omosome-scale sca olding o de no o genome assemblies based on ch oma in in e ac ions. Na u e bio echnology 31, 1119–1125 (2013). 11. Poland, J. A., B own, P. J., So ells, M. E. & Jannink, J.-L. De elopmen o high-densi y gene ic maps o ba ley and whea using a no el wo-enzyme geno yping-by-sequencing app oach. PLoS ONE 7, e32253 (2012). 12. Colmsee, C. e al. BARLEX— he Ba ley D a Genome Explo e . Mol Plan 8, 964–966 (2015). 13. Munoz-Ama iain, M. e al. Sequencing o 15 622 gene-bea ing BACs cla ifies he gene-dense egions o he ba ley genome. Plan Jou nal 84, 216–227 (2015). 14. Pasqua iello, M. e al. The ba ley F os esis ance-H2 locus. Func ional & in eg a i e genomics 14, 85–100 (2014). 15. Meye , M., S enzel, U. & Ho ei e , M. Pa allel agged sequencing on he 454 pla o m. Na u e p o ocols 3, 267–278 (2008). 16. S eue nagel, B. e al. De no o 454 sequencing o ba coded BAC pools o comp ehensi e gene su ey and genome analysis in he complex genome o ba ley. BMC genomics 10, 547 (2009). 17. Beie , S. e al. Mul iplex sequencing o bac e ial a ificial ch omosomes o assembling complex plan genomes. Plan bio- echnology jou nal 14, 1511–1522 (2016). 18. Samb ook, J. & Russell, D. W. Molecula cloning: a labo a o y manual. 3 d edi ion (Coldsp ing-Ha bou Labo a o y P ess, 2001). 19. Ai d, D. e al. Analyzing and minimizing PCR amplifica ion bias in Illumina sequencing lib a ies. Genome biology 12, R18 (2011). 20. Quail, M. A. e al. A la ge genome cen e ’s imp o emen s o he Illumina sequencing sys em. Na u e me hods 5, 1005–1010 (2008). 21. Asan e al. Pai ed-end sequencing o long- ange DNA agmen s o de no o assembly o la ge, complex Mammalian genomes by di ec in a-molecule liga ion. PLoS ONE 7, e46211 (2012). 22. Meye , M. & Ki che , M. Illumina sequencing lib a y p epa a ion o highly mul iplexed a ge cap u e and sequencing. Cold Sp ing Ha b P o oc 2010, pdb p o 5448 (2010). 23. Adey, A. e al. Rapid, low-inpu , low-bias cons uc ion o sho gun agmen lib a ies by high-densi y in i o ansposi ion. Genome biology 11, R119 (2010). 24. Lona di, S. e al. Combina o ial pooling enables selec i e sequencing o he ba ley gene space. PLoS compu a ional biology 9, e1003010 (2013). 25. Ze bino, D. R. & Bi ney, E. Vel e : algo i hms o de no o sho ead assembly using de B uijn g aphs. Genome esea ch 18, 821–829 (2008). 26. Ouni , R., Wanamake , S., Close, T. J. & Lona di, S. CLARK: as and accu a e classifica ion o me agenomic and genomic sequences using disc imina i e k-me s. BMC genomics 16, 236 (2015). 27. Zhang, Z., Schwa z, S., Wagne , L. & Mille , W. A g eedy algo i hm o aligning DNA sequences. Jou nal o compu a ional biology: a jou nal o compu a ional molecula cell biology 7, 203–214 (2000). 28. Che eux, B., We e , T. & Suhai, S. in Ge man con e ence on bioin o ma ics (1999); 45–56. 29. Taudien, S. e al. Sequencing o BAC pools by di e en nex gene a ion sequencing pla o ms and s a egies. BMC esea ch no es 4, 411 (2011). 30. Li, H. & Du bin, R. Fas and accu a e sho ead alignmen wi h Bu ows–Wheele ans o m. Bioin o ma ics 25, 1754–1760 (2009). 31. Boe ze , M., Henkel, C. V., Jansen, H. J., Bu le , D. & Pi o ano, W. Sca olding p e-assembled con igs using SSPACE. Bioin o ma ics 27, 578–579 (2011). 32. B enchley, R. e al. Analysis o he b ead whea genome using whole-genome sho gun sequencing. Na u e 491, 705–710 (2012). 33. And ews, S. Fas QC: a quali y con ol ool o high h oughpu sequence da a. A ailable online a : h p://www.bioin o ma ics. bab aham.ac.uk/p ojec s/ as qc (2010). 34. Legge , R. M., Rami ez-Gonzalez, R. H., Cla ijo, B. J., Wai e, D. & Da ey, R. P. Sequencing quali y assessmen ools o enable da a-d i en in o ma ics o high h oughpu genomics. F on ie s in gene ics 4, 288 (2013). 35. Simpson, J. T. e al. ABySS: a pa allel assemble o sho ead sequence da a. Genome esea ch 19, 1117–1123 (2009). 36. Magoc, T. & Salzbe g, S. L. FLASH: as leng h adjus men o sho eads o imp o e genome assemblies. Bioin o ma ics 27, 2957–2963 (2011). 37. Legge , R. M., Cla ijo, B. J., Clissold, L., Cla k, M. D. & Caccamo, M. Nex Clip: an analysis and ead p epa a ion ool o Nex e a Long Ma e Pai lib a ies. Bioin o ma ics 30, 566–568 (2014). 38. Luo, R. e al. SOAPdeno o2: an empi ically imp o ed memo y-e ficien sho - ead de no o assemble . GigaScience 1, 18 (2012). 39. Sla e , G. S. & Bi ney, E. Au oma ed gene a ion o heu is ics o biological sequence compa ison. BMC bioin o ma ics 6, 31 (2005). 40. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible sui e o u ili ies o compa ing genomic ea u es. Bioin o ma ics 26, 841–842 (2010). 41. R: A Language and En i onmen o S a is ical Compu ing (R Founda ion o S a is ical Compu ing, 2015). 42. Masche , M. e al. A ch omosome con o ma ion cap u e o de ed sequence o he ba ley genome. Na u e doi:10.1038/na u e22043 (2017). 43. Cao, H. e al. Rapid de ec ion o s uc u al a ia ion in a human genome using nanochannel-based genome mapping echnology. GigaScience 3, 1 (2014). 44. Chapman, J. A. e al. A whole-genome sho gun app oach o assembling and ancho ing he hexaploid b ead whea genome. Genome biology 16, 26 (2015). 45. Li, H. Aligning sequence eads, clone sequences and assembly con igs wi h BWA-MEM. P ep in a h ps://a xi .o g/pd / 1303.3997 2.pd (2013). 46. Masche , M., Wu, S., Amand, P. S., S ein, N. & Poland, J. Applica ion o geno yping-by-sequencing on semiconduc o sequencing pla o ms: a compa ison o gene ic and e e ence-based ma ke o de ing in ba ley. PLoS ONE 8, e76925 (2013). 47. Wu, Y., Bha , P. R., Close, T. J. & Lona di, S. E ficien and accu a e cons uc ion o gene ic linkage maps om he minimum spanning ee o a g aph. PLoS gene ics 4, e1000212 (2008). 48. Csa di, G. & Nepusz, T. The ig aph so wa e package o complex ne wo k esea ch, In e Jou nal, Complex Sys ems 1695 (2006). 49. P im, R. C. Sho es connec ion ne wo ks and some gene aliza ions. Bell sys em echnical jou nal 36, 1389–1401 (1957). 50. Wendle , N. e al. Unlocking he seconda y gene-pool o ba ley wi h nex -gene a ion sequencing. Plan bio echnology jou nal 12, 1122–1131 (2014). www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 21 51. Ma in, M. Cu adap emo es adap e sequences om high- h oughpu sequencing eads. EMBne . jou nal 17, 10–12 (2011). 52. Li, H. e al. The Sequence Alignmen /Map o ma and SAM ools. Bioin o ma ics 25, 2078–2079 (2009). 53. Ga ison, E. & Ma h, G. Haplo ype-based a ian de ec ion om sho - ead sequencing. P ep in a h ps://a xi .o g/pd / 1207.3907 2.pd (2012). 54. Kalho , R., Tjong, H., Jaya hilaka, N., Albe , F. & Chen, L. Genome a chi ec u es e ealed by e he ed ch omosome con o ma ion cap u e and popula ion-based modeling. Na u e bio echnology 30, 90–98 (2012). 55. Ma sumo o, T. e al. Comp ehensi e sequence analysis o 24,783 ba ley ull-leng h cDNAs de i ed om 12 clone lib a ies. Plan physiology 156, 20–28 (2011). 56. Wu, T. D. & Wa anabe, C. K. GMAP: a genomic mapping and alignmen p og am o mRNA and EST sequences. Bioin o ma ics 21, 1859–1875 (2005). 57. DeP is o, M. A. e al. A amewo k o a ia ion disco e y and geno yping using nex -gene a ion DNA sequencing da a. Na u e gene ics 43, 491–498 (2011). 58. A end, D. e al. PGP eposi o y: a plan phenomics and genomics da a publica ion in as uc u e. Da abase 2016, baw033 (2016). 59. A end, D. e al. e!DAL--a amewo k o s o e, sha e and publish esea ch da a. BMC bioin o ma ics 15, 214 (2014). 60. Künzel, G., Ko zun, L. & Meis e , A. Cy ologically in eg a ed physical es ic ion agmen leng h polymo phism maps o he ba ley genome based on ansloca ion b eakpoin s. Gene ics 154, 397–412 (2000). 61. Aliye a-Schno , L. e al. Cy ogene ic mapping wi h cen ome ic bac e ial a ificial ch omosomes con igs shows ha his ecombina ion-poo egion comp ises mo e han hal o ba ley ch omosome 3H. The Plan Jou nal 84, 385–394 (2015). Da a Ci a ions 1. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9062 (2016). 2. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9097 (2016). 3. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9098 (2016). 4. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9099 (2016). 5. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9100 (2016). 6. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9101 (2016). 7. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9102 (2016). 8. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9103 (2016). 9. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9104 (2016). 10. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8576 (2016). 11. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8577 (2016). 12. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8578 (2016). 13. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9619 (2016). 14. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8579 (2016). 15. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8580 (2016). 16. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9429 (2016). 17. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9430 (2016). 18. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9431 (2016). 19. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB10963 (2016). 20. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11489 (2016). 21. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB12096 (2016). 22. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11758 (2016). 23. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9428 (2016). 24. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11991 (2016). 25. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9427 (2016). 26. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11798 (2016). 27. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11992 (2016). 28. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB13020 (2016). 29. Muñoz-Ama iaín, M. e al. NCBI BioP ojec PRJNA198204 (2015). 30. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/21 (2016). 31. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/28 (2016). 32. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/12 (2016). 33. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/31 (2016). 34. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB13028 (2016). 35. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/33 (2016). 36. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/22 (2016). 37. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/30 (2016). 38. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB14130 (2016). 39. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/29 (2016). 40. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB14169 (2016). 41. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/20 (2016). 42. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/34 (2016). 43. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/27 (2016). 44. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/36 (2016). 45. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/23 (2016). 46. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/24 (2016). 47. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/25 (2016). 48. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/26 (2016). 49. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/35 (2016). 50. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/37 (2016). 51. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/17 (2016). 52. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/19 (2016). Acknowledgemen s This wo k was ca ied ou unde he auspices o he In e na ional Ba ley Genome Sequencing Conso ium and suppo ed om he ollowing unding sou ces: Ge man Minis y o Educa ion and Resea ch (BMBF) g an 0314000 ‘BARLEX’and 0315954 ‘TRITEX’ o M.P., U.S. and N.S and 031A536 www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 22 ‘de.NBI’ o U.S. Leibniz Associa ion g an (‘Pak . Fo schung und Inno a ion’)‘sequencing ba ley ch omosome 3H’ o N.S. and U.S.; Sco ish Go e nmen /UK Bio echnology and Biological Sciences Resea ch Council (BBSRC) g an BB/100663X/1 o R.W, P.E.H., J.R.; BBSRC g an s BB/I008357/1 o M. D.C., M.C. and BB/I008071/1 o P.K.; o Finland g an 266430 and a BioNano g an o A.H.S.; Ca lsbe g Founda ion g an n . 2012_01_0461 o he Ca lsbe g Resea ch Labo a o y; G ain Resea ch and De elopmen Co po a ion (GRDC) g an DAW00233 o C.L. and P.L.; Depa men o Ag icul u al and Food, Go e nmen o Wes e n Aus alia g an 681 o C.L.; Na ional Na u al Science Founda ion o China (NSFC) g an 31129005 o C.L. and G.Zhang; NSFC g an 31330055 o G.Zhang.; Czech Minis y o Educa ion, You h and Spo s g an LO1204 o J.D.; Na ional Science Founda ion g an DBI 0321756 ‘Coupling EST and Bac e ial A ificial Ch omosome Resou ces o Access he Ba ley Genome’ o T.J.C. and S.L.; Uni ed S a es Depa men o Ag icul u e (USDA), Ag icul u e and Food Resea ch Ini ia i e Plan Genome, Gene ics and B eeding P og am o USDA-CSREES-NIFA g an 2009-65300-05645 ‘Ad ancing he Ba ley Genome’and 2011-68002-30029 ‘T i iceaeCAP’ o T.J.C., S.L. and G.J.M.; Uni ed S a es Na ional Science Founda ion (NSF)-ABI g an DBI-1062301 o T.J.C. and S.L.; Uni e si y o Cali o nia g an CA-R-BPS-5306-H o T.J.C and S.L.;Na ional Science Founda ion g an DBI 0321756 ‘Algo i hms o Genome Assembly o Ul a-deep Sequencing Da a’ o S.L. Nex -gene a ion sequencing and lib a y cons uc ion was deli e ed ia he BBSRC Na ional Capabili y in Genomics (BB/J010375/1) a Ea lham Ins i u e ( o me ly The Genome Analysis Cen e) by membe s o he Pla o ms and Pipelines g oup and BBSRC Ins i u e S a egic P og amme unding o Bioin o ma ics (BB/J004669/1) o M.D.C., S.A. and M.C. We g a e ully acknowledge: (1) he excellen echnical assis ance by Susanne König, Manuela Knau , Uli Beie , Anne Kusse ow, Ka in T nka, Ines Walde, Sand a D iesslein, Cyn hia Voss; (2) Do een S engel, Anne Fiebig, Thomas Münch, Danu a Schüle and Daniel A end and Ma hias Lange o sequence aw da a managemen and da a submission o EMBL/ENA and egis a ion o DOIs; (3) D Hélène Be ges, A naud Bellec and Sonia Vau in (CNRGV) o managemen and dis ibu ion o ba ley BAC lib a ies; (4) And eas G ane and Da id Ma shall o scien ific discussions. Au ho Con ibu ions BAC sequencing and assembly (1H, 3H, 4H): S.B., A.Himmelbach, S.T., M.F., M.G., M.M., U.S. (co-leade ), M.P. (co-leade ), N.S. (leade ); BAC sequencing and assembly (2H, unassigned): D.S., D.H., S. A. (co-leade ), M.D.C. (co-leade ), M.C. (co-leade ), R.W. (leade ); BAC sequencing and assembly (5H, 7H): X.Z., R.A.B., Q.Z., C.T., J.K.M., B.C., G.Zhou, F.D., Y.H., S.Y., S.Cao, S.Wang, X.L., M.I.B., P.L., G.Zhang (co-leade ), C.Li (leade ); BAC sequencing and assembly (6H): S.B., S.Wang, C.Lin, H.L., U.S., M. H. (co-leade ), I.B. (leade ); BAC sequencing (gene-bea ing): M.M.-A., R.O., S.Wanamake , S.L. (co-leade ), T.J.C. (leade ); Op ical mapping: A.Has ie, H.Š., J.T., H.S., J.V., S.Chan, M.M., N.S., J.D., A.H.S. (leade ); Ch omosome con o ma ion cap u e: A.Himmelbach, S.G., M.M. (co-leade ), N.S. (leade ); Pseudomolecule cons uc ion: M.M. (leade ), S.B., C.C., D.B., T.S., P.K., N.S., U.S. (co-leade ); Valida ion: L.L., M.B., L.A.-S., A.Houben, J.A.P., N.S., G.J.M., M.M. (leade ). All au ho s ead and commen ed on he manusc ip . Addi ional in o ma ion Compe ing financial in e es s: The au ho s decla e no compe ing financial in e es s. How o ci e his a icle: Beie , S. e al. Cons uc ion o a map-based e e ence genome sequence o ba ley, Ho deum ulga e L. Sci. Da a 4:170044 doi: 10.1038/sda a.2017.44 (2017). Publishe ’s no e: Sp inge Na u e emains neu al wi h ega d o ju isdic ional claims in published maps and ins i u ional a filia ions. This wo k is licensed unde a C ea i e Commons A ibu ion 4.0 In e na ional License. The images o o he hi d pa y ma e ial in his a icle a e included in he a icle’s C ea i e Commons license, unless indica ed o he wise in he c edi line; i he ma e ial is no included unde he C ea i e Commons license, use s will need o ob ain pe mission om he license holde o ep oduce he ma e ial. To iew a copy o his license, isi h p://c ea i ecommons.o g/licenses/by/4.0 Me ada a associa ed wi h his Da a Desc ip o is a ailable a h p://www.na u e.com/sda a/ and is eleased unde he CC0 wai e o maximize euse. © The Au ho (s) 2017 www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 23 Sebas ian Beie 1,* , Axel Himmelbach 1,* , Ch is ian Colmsee 1 , Xiao-Qi Zhang 2 , Robe o A. Ba e o 3 , Qisen Zhang 4 , Lin Li 5 , Micha Baye 6 , Daniel Bolse 7 , S e an Taudien 8 , Ma co G o h 8 , Ma ius Felde 8 , Alex Has ie 9 , Hana Šimko á 10 , Helena S aňko á 10 , Jan V ána 10 , Saki Chan 9 , Ma ía Muñoz-Ama iaín 11 , Rachid Ouni 12 , S e e Wanamake 11 , Thomas Schmu ze 1 , Lala Aliye a-Schno 1 , S e ano G asso 13 , Jaakko Tanskanen 14 , Dha anya Sampa h 15 , Da en Hea ens 15 , Sujie Cao 16 , B e Chapman 3 , Fei Dai 17 , Yong Han 17 ,HuaLi 16 , Xuan Li 16 , Chongyun Lin 16 , John K. McCooke 3 , Cong Tan 3 , Songbo Wang 16 , Shuya Yin 17 , Gao eng Zhou 2 , Jesse A. Poland 18 , Ma hew I. Bellga d 3 , And eas Houben 1 , Ja osla Doležel 10 , Sa ah Ayling 15 , S e ano Lona di 12 , Pe e Lang idge 19 , Ga y J. Muehlbaue 5,20 , Paul Ke sey 7 , Ma hew D. Cla k 15,21 , Ma io Caccamo 15,22 , Alan H. Schulman 14 , Ma hias Pla ze 8 , Timo hy J. Close 11 , Ma s Hansson 23 , Guoping Zhang 17 , Ilka B aumann 24 , Chengdao Li 2,25,26 , Robbie Waugh 6,27 , Uwe Scholz 1 , Nils S ein 1,28 & Ma in Masche 1,29 1 Leibniz Ins i u e o Plan Gene ics and C op Plan Resea ch (IPK) Ga e sleben, 06466 Seeland, Ge many. 2 School o Ve e ina y and Li e Sciences, Mu doch Uni e si y, Mu doch, Wes e n Aus alia 6150, Aus alia. 3 Cen e o Compa a i e Genomics, Mu doch Uni e si y, Mu doch, Wes e n Aus alia 6150, Aus alia. 4 Aus alian Expo G ains Inno a ion Cen e, Sou h Pe h, Wes e n Aus alia 6151, Aus alia. 5 Depa men o Ag onomy and Plan Gene ics, Uni e si y o Minneso a, S Paul, Minneso a 55108, USA. 6 The James Hu on Ins i u e, Dundee DD2 5DA, UK. 7 Eu opean Molecula Biology Labo a o y—The Eu opean Bioin o ma ics Ins i u e, Hinx on CB10 1SD, UK. 8 Leibniz Ins i u e on Aging—F i z Lipmann Ins i u e (FLI), 07745 Jena, Ge many. 9 BioNano Genomics Inc., San Diego, Cali o nia 92121, USA. 10 Ins i u e o Expe imen al Bo any, Cen e o he Region Haná o Bio echnological and Ag icul u al Resea ch, 78371 Olomouc, Czech Republic. 11 Depa men o Bo any & Plan Sciences, Uni e si y o Cali o nia, Ri e side, Ri e side, Cali o nia 92521, USA. 12 Depa men o Compu e Science and Enginee ing, Uni e si y o Cali o nia, Ri e side, Ri e side, Cali o nia 92521, USA. 13 Depa men o Ag icul u al and En i onmen al Sciences, Uni e si y o Udine, 33100 Udine, I aly. 14 G een Technology, Na u al Resou ces Ins i u e (Luke), Viikki Plan Science Cen e, and Ins i u e o Bio echnology, Uni e si y o Helsinki, 00014 Helsinki, Finland. 15 Ea lham Ins i u e, No wich NR4 7UH, UK. 16 BGI-Shenzhen, Shenzhen 518083, China. 17 College o Ag icul u e and Bio echnology, Zhejiang Uni e si y, Hangzhou 310058, China. 18 Kansas S a e Uni e si y, Whea Gene ics Resou ce Cen e , Depa men o Plan Pa hology and Depa men o Ag onomy, Manha an, Kansas 66506, USA. 19 School o Ag icul u e, Uni e si y o Adelaide, U b ae, Sou h Aus alia 5064, Aus alia. 20 Depa men o Plan and Mic obial Biology, Uni e si y o Minneso a, S Paul, Minneso a 55108, USA. 21 School o En i onmen al Sciences, Uni e si y o Eas Anglia, No wich NR4 7UH, UK. 22 Na ional Ins i u e o Ag icul u al Bo any, Camb idge CB3 0LE, UK. 23 Depa men o Biology, Lund Uni e si y, 22362 Lund, Sweden. 24 Ca lsbe g Resea ch Labo a o y, 1799 Copenhagen, Denma k. 25 Depa men o Ag icul u e and Food, Go e nmen o Wes e n Aus alia, Sou h Pe h, Wes e n Aus alia 6150, Aus alia. 26 Hubei Collabo a i e Inno a ion Cen e o G ain Indus y, Yang ze Uni e si y, Jingzhou, Hubei 434025, China. 27 School o Li e Sciences, Uni e si y o Dundee, Dundee DD2 5DA, UK. 28 School o Plan Biology, Uni e si y o Wes e n Aus alia, C awley 6009, Aus alia. 29 Ge man Cen e o In eg a i e Biodi e si y Resea ch (iDi ) Halle-Jena-Leipzig, 04103 Leipzig, Ge many. *These au ho s con ibu ed equally o his wo k. www.na u e.com/sda a/ SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 24