Da a Desc ip o : Cons uc ion o
a map-based e e ence genome
sequence o ba ley, Ho deum
ulga e L.
Sebas ian Beie e al.
#
Ba ley (Ho deum ulga e L.) is a ce eal g ass mainly used as animal odde and aw ma e ial o he mal ing
indus y. The map-based e e ence genome sequence o ba ley c . ‘Mo ex’was cons uc ed by he
In e na ional Ba ley Genome Sequencing Conso ium (IBSC) using hie a chical sho gun sequencing. He e,
we epo he expe imen al and compu a ional p ocedu es o (i) sequence and assemble mo e han 80,000
bac e ial a ificial ch omosome (BAC) clones along he minimum iling pa h o a genome-wide physical
map, (ii) find and alida e o e laps be ween adjacen BACs, (iii) cons uc 4,265 non- edundan sequence
sca olds ep esen ing clus e s o o e lapping BACs, and (i ) o de and o ien hese BAC clus e s along he
se en ba ley ch omosomes using posi ional in o ma ion p o ided by dense gene ic maps, an op ical map
and ch omosome con o ma ion cap u e sequencing (Hi-C). In eg a i e access o hese sequence and
mapping esou ces is p o ided by he ba ley genome explo e (BARLEX).
Design Type(s) genome assembly
Measu emen Type(s) whole genome sequencing assay
Technology Type(s) DNA sequencing
Fac o Type(s) lib a y p epa a ion
Sample Cha ac e is ic(s) Ho deum ulga e
Co espondence and eques s o ma e ials should be add essed o M.M. (email: masche @ipk-ga e sleben.de).
#A ull lis o au ho s and hei a filia ions appea s a he end o he pape .
OPEN
Recei ed: 26 Augus 2016
Accep ed: 9Feb ua y 2017
Published: 27 Ap il 2017
www.na u e.com/scien i icda a
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 1
Backg ound & Summa y
Ba ley (Ho deum ulga e L.) is a ce eal g ass o g ea ag onomical impo ance. The goal o he
In e na ional Ba ley Genome Sequencing Conso ium (IBSC) is he cons uc ion o a map-based
e e ence sequence assembly o ba ley cul i a ‘Mo ex’by means o hie a chical sho gun sequencing
1
.
Towa ds his aim, he ba ley genomics communi y has de eloped an a ay o genome-wide physical and
gene ic mapping esou ces. These include lib a ies o bac e ial a ificial ch omosomes (BACs)
2
,
a genome-wide physical map
3
, a d a whole genome sho gun (WGS) assembly
4
and an ul a-dense
gene ic map
5
. The las s age on he oad owa ds he e e ence genome is he sho gun sequencing o BAC
clones along a minimum iling pa h o he genome defined by he physical map. The ad ances in high-
h oughpu sequencing echnology enabled his ask o be comple ed in a much sho e ime ame han
was equi ed o he comple ion o , o ins ance, he human
6
and maize
7
genomes. In addi ion o he
gene a ion o BAC aw sequence da a, we cons uc ed (i) physical genome maps by single-molecule
op ical mapping in nanochannels
8
and by ch omosome con o ma ion cap u e sequencing (Hi-C)
9,10
, and
(ii) a high- esolu ion gene ic map o a la ge bi-pa en al mapping popula ion h ough geno yping-by-
sequencing
11
. We unde ook he sequence assembly o indi idual BACs, he cons uc ion o la ge
sequence sca olds by me ging sequences om adjacen clones and he in eg a ion o hese supe -
sca olds wi h he a ious genome-wide mapping esou ces cons uc ed in he p esen e o as well as
hose published p e iously
3,5
. The final ou come o his app oach was he cons uc ion o
‘pseudomolecules’, i.e., con iguous sequence sca olds ep esen ing he se en ch omosomes o ba ley.
We ha e submi ed he ele an aw da a o public sequence da a a chi es, made analysis esul s
a ailable unde pe manen digi al objec iden ifie s (DOIs) and en e ed he posi ional in o ma ion used
o pseudomolecule cons uc ion in o a bespoke in o ma ion managemen sys em, he BARLEX genome
explo e
12
. He e, we gi e (i) a comp ehensi e o e iew o da ase s used o assembling he ba ley genome
and me hods employed in hei gene a ion, (ii) a de ailed desc ip ion o we -lab p ocedu es o BAC
sequencing and he bioin o ma ics wo kflow o he sequence assembly and da a in eg a ion p ocedu es
oge he wi h an ou line o (iii) hei b owsable p esen a ion in an online da abase. These esou ces
documen he cons uc ion o he map-based e e ence sequence o he ba ley genome and will enable
esea che s o inspec he e idence used o assemble, o de and o ien sequence sca olds and may guide
he u he imp o emen o he genome sequence wi h complemen a y da a se s.
Me hods
The main s eps o he cons uc ion o he map-based e e ence sequence o he ba ley genome we e
(i) sho gun and ma e-pai sequencing o BAC clones, (ii) sequence assembly o indi idual BAC clones
and (iii) he cons uc ion o a pseudomolecule sequences by me ging he sequences o adjacen BACs in o
supe -sca olds and o de ing hese using a ious sou ces o posi ional in o ma ion such as physical maps,
op ical map and ch omosome con o ma ion cap u e. A schema ic o e iew o ou expe imen al
p ocedu es is gi en in Fig. 1.
BAC sequencing
Iden ifica ion and analysis o gene-con aining BACs. Isola ion o gene-con aining BACs,
cons uc ion o a minimal iling pa h (MTP), sequencing o MTP clones and he anno a ion o genes
we e essen ially as desc ibed p e iously
13
.
Sho gun and ma e-pai sequencing o MTP-BACs. Sequencing o MTP-BACs was conduc ed in ou
labo a o ies (Leibniz Ins i u e on Aging—F i z Lipmann Ins i u e (FLI) Jena, Leibniz Ins i u e o Plan
Gene ics and C op Plan Resea ch (IPK) Ga e sleben, Beijing Genomics Ins i u e (BGI) and Ea lham
Ins i u e (EI) No wich). Depending on he ins umen a ion and es ablished p o ocols, cus omized
app oaches we e aken o sequence he ba ley MTP BACs.
Ba ley ch omosomes 1H, 3H and 4H (IPK and FLI)
Sho gun sequencing o MTP BACs
Du ing he ini ial phase, BACs mos ly om ch omosome 3H (4870 clones) and a small
numbe o clones om o he ch omosomes (34 om 1H; 31 om 2H; 50 om 4H; 101 om 5H;
33 om 6H; 64 om 7H; 107 om ‘0H’) we e sho gun sequenced using he Roche/454 GS FLX
de ice (Da a Ci a ion 1, Da a Ci a ion 2, Da a Ci a ion 3, Da a Ci a ion 4, Da a Ci a ion 5,
Da a Ci a ion 6, Da a Ci a ion 7, Da a Ci a ion 8, Da a Ci a ion 9). BAC DNA was p epa ed using
a modified alkaline lysis p o ocol
14
. Cons uc ion o ba coded 454 sequencing lib a ies and sequencing using
he Roche pla o m we e pe o med as desc ibed
15,16
. The emaining BAC clones om ch omosomes 1H,
3H and 4H we e sho gun sequenced employing Illumina ins umen s. BAC DNA isola ion, lib a y
cons uc ion, sequencing-by-syn hesis (pai ed-end, 2 × 100 cycles) using he Illumina HiSeq2000 de ice was
pe o med as desc ibed
17
(Da a Ci a ion 10, Da a Ci a ion 11, Da a Ci a ion 12, Da a Ci a ion 13). Pools o
up o 667 BACs we e indi idually ba coded and sequenced on one HiSeq2000 lane.
In addi ion, he Illumina GAIIx, HiSeq2500 and MiSeq machines we e u ilized o sequence pools o up
o 384 clones pe lane as desc ibed p e iously
17
.
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 2
Ma e-pai sequencing o MTP BACs
Fo sca olding o ch omosomes 1H, 3H and 4H s anda d Illumina Nex e a ma e-pai lib a ies
(span size: 8 kb) o BAC pools up o 384 BACs we e cons uc ed and sequenced using he
Illumina HiSeq2000 (pai ed end, 2 × 100 cycles) and MiSeq (pai ed end, 2 × 250 cycles) as desc ibed
17
(Da a Ci a ion 14, Da a Ci a ion 15).
Ba ley ch omosomes 5H, 6H and 7H (BGI)
Sho gun sequencing o MTP BACs
Bac e ial s a e cul u es we e inocula ed in 0.4 ml 2 × YT liquid medium
18
supplemen ed wi h
chlo amphenicol (17.5 μgml
1
) in 2 ml polyp opylene 96-deep well-pla es sealed wi h gas-pe meable oil
and incuba ed a 37 °C o 14 h in a shaking incuba o (210 .p.m.). Fo DNA isola ion duplica es o
cul u es (1 ml 2 × YT liquid medium con aining 17.5 μgml
1
chlo amphenicol) we e inocula ed wi h
50 μl s a e cul u e and incuba ed (37 °C, 14 h, 210 .p.m.). BAC DNA was isola ed using he alkaline
lysis me hod essen ially as desc ibed p e iously
17
. The DNA was dissol ed (o e nigh , 4 °C) in 64 μlTE
(pH 8.0) con aining RNase A (30 μgml
1
) and s o ed a −20 °C. BAC plasmid DNA (0.5–2.0 μgin60μl)
was andomly agmen ed by ocused-ul asonica o (Co a is LE220 ins umen : 21% du y ac o ,
500 PIP, 500 cycles pe bu s , 70 s ea men ime) in 96-well pla es (Axygen, PCR-96M2-HS-C) o
an a e age size o 250–750 bp. The DNA agmen s we e pu ified using magne ic beads
(GeneOn Pu ifica ion ki , GO-PCRC-5000) acco ding o he manu ac u e ’s ins uc ions. DNA was
p ecipi a ed by adding 10 μl magne ic bead suspension and 75 μl Binding Bu e . The samples we e mixed
and incuba ed a oom empe a u e o 5 min. Beads con aining he DNA we e eclaimed by using
a magne (96S Supe Magne Pla e, ALPAQUA, A001322), and he clea supe na an was disca ded. The
beads we e washed wice wi h 200 μl o 70% e hanol and d ied comple ely. Fo he elu ion o DNA he
beads we e suspended in 42 μl Elu ion Bu e (EB, 10 mM T is-Cl, pH 8.5) and incuba ed (5 min). The
pla e was placed on he magne , and he supe na an (40 μl) was ans e ed in o new 96-well pla es.
End- epai and A-Tailing we e pe o med as desc ibed
19
. The eac ion clean-ups we e pe o med wi h
GeneOn magne ic beads as desc ibed abo e. Ba code adap e s (1 μl, 20 μM) o he fi s index
we e liga ed o he s icky ends o DNA agmen s by using T4 DNA ligase
19
, incuba ed a 16 °C o a
leas 12 h. Each indi idual sample was p o ided wi h a di e en ba code o a se o 384 di e en indices
(adap e and ba code sequences a e a ailable upon eques ). Equal olumes o he 384 indi idually
ba coded adap e -liga ed p oduc s we e pooled. The pooled DNA was p ecipi a ed by adding
20 μl GeneOn magne ic beads and 650 μl Binding Bu e (GeneOn Pu ifica ion Ki , GO-PCRC-5000)
BAC DNA
p epa a ion
Pai ed-end
lib a y
cons uc ion
Ma e-pai
lib a y
cons uc ion
Quan i ica ion,
pooling and size
ac iona ion
Sequencing-by-
syn hesis
(Illumina)
Remo al o low
quali y
sequences &
con amina ion
Indi idual
BAC
assembly
Quan i ica ion,
pooling and size
ac iona ion
Sequencing-by-
syn hesis
(Illumina)
Remo al o low
quali y
sequences &
con amina ion
Mapping &
indi idual BAC
sca olding
Indi idual
BAC
sca olds
BAC
sca olds
FPC / BES
da a
POPSEQ
map
Bionano
map da a
Con o ma ion
cap u e da a
(HiC / TCC)
BAC o e lap
clus e s
BLAST
analysis
Non-
edundan
sequence
Con o ma ion
cap u e map
(HiC map)
AGP gene a ion &
Pseudomolecule
sequence
Figu e 1. Assembly wo kflow. (a) Assembly o indi idual BAC clones om pai ed-end and ma e-pai ead
da a. (b) Da a in eg a ion p ocedu es o pseudomolecule cons uc ion.
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 3
o 500 μl pooled DNA. The suspension was mixed and incuba ed a oom empe a u e o 5 min. The
beads con aining he DNA we e eclaimed using a magne , and he clea supe na an was disca ded. The
beads we e washed wice wi h 500 μl o 70% e hanol and d ied comple ely. The DNA was elu ed in
52 μl EB. The sample was size-sepa a ed by using s anda d aga ose gel elec opho esis (2% aga ose gel,
HyAga ose, 16250). DNA was e ealed using e hidium b omide and exci a ion by isible blue ligh
emi ed om a Da k Reade blue ligh ansillumina o (Cla e Chemical Resea ch) o selec he a ge
agmen s (580–620 bp). The a ge egion was ex ac ed in 27 μl EB using he QIAquick Gel Ex ac ion
ki (QIAGEN). The second index was in oduced using he adap e -liga ed p oduc s as empla e DNA
(98 °C o 30 s, 10 cycles o : 98 °C o 10 s, 65 °C o 30 s and 72 °C o 30 s, final ex ension 72 °C o
5 min) (Enzyma ics, CM0075) and PCR p oduc s ( a ge egion: 580–620 bp) we e eco e ed by aga ose
gel elec opho esis (2% aga ose gel, HyAga ose, 16250) as desc ibed abo e. Index p ime s we e used
o ba coding each 384 pooled BAC samples (index p ime sequences a e a ailable upon eques ). The
a e age size o he PCR p oduc s was de e mined by using an Agilen 2100 Bioanalyze (Agilen DNA
1,000 Reagen s). Typical a e age size o he lib a ies was be ween 574 o 674 bp. PCR p oduc s we e
quan ified using eal- ime PCR and pooled o sequencing in equal p opo ion
20
. Pai ed-end sequencing
(2 × 100 cycles; fi s index: 11 cycles, second index: 8 cycles) was pe o med on he Illumina HiSeq2000
pla o m (Da a Ci a ion 16, Da a Ci a ion 17, Da a Ci a ion 18).
Ma e-pai sequencing o MTP BACs
Fo he cons uc ion o ma e-pai lib a ies (10 and 20 kb span size), 96 BACs co esponding o 6 μg DNA
we e pooled in o one ube. The DNA was agmen ed o 10 o 20 kb by using he Hyd oShea DNA
Shea ing sys em om GeneMachines (10 kb: la ge assembly, speed code 12, cycles 12, olume 250 μl;
20 kb: la ge assembly, speed code 13, cycles 20, olume 150 μl). Following DNA agmen a ion, he
agmen s we e pu ified by using 0.6 olumes magne ic beads (Axygen, MAG-PCR-CL-250). The samples
we e mixed and incuba ed a oom empe a u e o 10 min. Beads con aining he DNA we e eclaimed by
using a magne pla e (96S Supe Magne Pla e, ALPAQUA, A001322), and he clea supe na an was
disca ded. The beads we e washed wice wi h 500 μl o 70% e hanol and d ied comple ely. Fo he elu ion
o DNA he beads we e esuspended in 80 μl EB. End- epai and bio in-labeling we e pe o med as
desc ibed
21
. End- epai ed DNA was pu ified using 0.6 olumes magne ic beads (Axygen, MAG-PCR-
CL-250) as desc ibed o he pu ifica ion o hyd o-shea ed DNA. The DNA was elu ed in 79 μl EB. 20 kb
lib a ies (20–26 kb ange) we e size-selec ed using aga ose gel (0.6%) elec opho esis. The liga ion o he
lib a ies, was pe o med by adding 1 μl Ba code Adap o (20 μM, sequences a e a ailable upon eques ),
10 μl T4 DNA ligase (Enzyma ics, L603-HC) in a o al olume o 100 μl (20 °C, 15 min). 15 indi idually
ba coded adap o -liga ed DNAs (10 kb) we e pooled in equimola manne and size- ac iona ed
(9–11 kb) using aga ose gel (0.6%) elec opho esis. DNA ci cula iza ion and emo al o non-ci cula ized
DNA was as desc ibed
21
. The DNA was isola ed om he gel using he QIAquick Gel Ex ac ion ki as
desc ibed by he manu ac u e (QIAGEN). Ci cula DNA was agmen ed using he Co a is S2 de ice
(10% du y cycle, 10 in ensi y, 1,000 bu s s pe second, 22 min (11 min) ea men ime o 10 kb (20 kb)
lib a ies in TC13 Co a is ubes), and bio inyla ed agmen s de i ed om ue ma e-pai liga ion e en s
we e pu ified using s ep a idin-coupled Dynabeads (M-280, In i ogen)
19
. Ends o he DNA agmen s
we e epai ed and p o ided wi h Illumina pai ed-end adap e s as desc ibed o he cons uc ion o
sho gun lib a ies. The bead-bound DNA was PCR-amplified using Phusion polyme ase (NEB) (98 °C o
30 s, 18 cycles o : 98 °C o 10 s, 65 °C o 30 s, 72 °C o 30 s and a final ex ension: 72 °C o 5 min) using
manu ac u e ’s p o ocols (NEB). Size-selec ion was essen ially pe o med as desc ibed o sho gun lib a y
cons uc ion. Fo he 10 kb (20 kb) ma e-pai lib a ies, DNA in he size ange be ween 270–420 bp
(400–600 bp) was isola ed and pu ified using he QIAquick Gel Ex ac ion ki acco ding o
manu ac u e ’s ins uc ions (QIAGEN). The a e age size o he pai ed-end BAC lib a ies was de e mined
elec opho e ically using an Agilen 2100 Bioanalyze (Agilen DNA 1,000 Reagen s). Lib a ies we e
quan ified using Real-Time PCR
20
. The ma e-pai lib a ies we e pai ed-end sequenced using he Illumina
HiSeq2500 de ice (10 kb lib a y: 150 cycles, 20 kb ma e-pai lib a y 50 cycles). Raw da a a e a ailable as
Da a Ci a ion 19, Da a Ci a ion 20, Da a Ci a ion 21).
Ba ley ch omosomes 2H and 0H (EI)
Sho gun sequencing o MTP BACs
QRep 384 Pin Replica o s (Molecula De ices, New Mol on, UK) we e used o inocula e clones om
s ock pla es in o 384 squa e deep well cul u e pla es con aining 140 μl 2 × YT media supplemen ed wi h
12.5 μgml
1
chlo amphenicol
18
. The cul u e pla es we e sealed wi h a gas pe meable seal and incuba ed
o 22 h a 37 °C in a shaking incuba o (200 .p.m.). Cells we e ha es ed by cen i uga ion (20 min,
3,220 g, 4 °C), he supe na an was disca ded. BAC DNA was p epa ed using a modified alkaline lysis
p o ocol (Beckman Coul ie , High Wycombe, UK). Cell pelle s we e esuspended in 8 μl o Resuspension
Bu e (RE1) using a Mic opla e Shake TiMix 5 con ol (Edmund-Buehle , Hechingen, Ge many)
(10 min, 1,400 .p.m.). Cells we e lysed by adding 8 μl o he lysis solu ion (L2). A e shaking (5 min,
500 .p.m.) 8 μl o cold Neu alisa ion Bu e (N3) we e added. The pla e was shaken (10 min, 500 .p.m.)
ollowed by a cen i uga ion (20 min, 3,220 g, 4 °C). The clea supe na an (14.33 μl) was ans e ed
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 4
o a 384 well PCR pla e, which con ained 1 μl o CosMc beads pe well. The pla e was mixed b iefly
(500 .p.m.), 10 μl o isop opanol was added and he suspension was mixed b iefly again (500 .p.m.). The
pla e was incuba ed a oom empe a u e o 15 min o allow p ecipi a ion o he DNA on o he beads.
The pla e con aining he DNA p ecipi a e was mo ed on o a 96 pin 384 well pla e compa ible magne
(Alpaqua, Be e ley, MA, USA) and le o 5 min o he beads o pelle . The supe na an was disca ded
and he beads we e washed h ee imes wi h 20 μl 70% e hanol while placed in he magne and ai d ied
( oom empe a u e, 5 min). The DNA was elu ed om he beads in 20 μl o 10 mM T is HCl (pH 8.0)
and ans e ed o a esh 384 well PCR pla e. To emo e con amina ing hos E. coli gDNA samples we e
ea ed wi h Epicen e Plasmid Sa e ATP dependen DNase (Cambio, Camb idge, UK), which diges s he
agmen ed E. coli and nicked BAC DNA bu lea es supe coiled BAC DNA in ac . To 20 μl o DNA 2.5 μl
o 10x Reac ion bu e , 1 μl 25 mM ATP, 0.1 μl ATP dependen DNase (10 u μl
1
) and 1.4 μl wa e was
added, and he samples we e incuba ed a 37 °C (8 h) ollowed by 70 °C (20 min) o inac i a e he DNase.
Sequencing lib a ies (single index) om he ini ial six een 384 well pla es o BACs (2H ch omosome)
we e cons uc ed in 384 well PCR pla es (Fo i ude, Wo on, UK) using he Epicen e Nex e a Ki
(Epicen e, Madison, WI, USA) and Robus 2G Taq polyme ase (Kapa Biosciences, London, UK). The
384 adap e oligos wi h 9 bp ba codes each wi h a hamming dis ance o 4 (adap e sequences a e a ailable
upon eques ) we e designed using s anda d guidelines
22
. B iefly, 1 μl o BAC DNA, 1 μl Nex e a HMW
5 × Reac ion Bu e , 1 μl o Nex e a Enzyme (dilu ed 50- old in 50% glyce ol, 0.5 × TE pH 8.0) and 2 μlo
wa e we e combined and incuba ed (5 min, 55 °C) as desc ibed
23
. Fo he dena u a ion o he Tn5
polyme ase, 15 μl PB Bu e (Qiagen, Manches e , UK) and o he eac ion clean-up, 20 μl AMPu e XP
(Beckman, High Wycombe, UK) beads we e added using a Calipe Sciclone Robo (Pe kin Elme ,
Co en y, UK). Following an incuba ion (5 min, oom empe a u e), he p ecipi a ed agmen ed DNA
was pu ified using a 96 well ing Magne (Alpaqua, Be e ly, MA, USA). The beads we e washed wice
wi h 20 μl 70% e hanol while placed in he magne be o e being ai d ied o 5 min. The agmen ed DNA
was elu ed in 5 μl 10 mM T is HCl, pH 8.0 and ans e ed o a esh 384 well PCR pla e. To 5 μl pu ified,
agmen ed DNA 2 μl o 5 × 2G B Reac ion bu e , 0.2 μl o 10 mM dNTPs, 0.1 μl o Robus 2G Taq
polyme ase, 0.2 μl o 50 × Nex e a P ime Cock ail and 2.5 μl 0.2 μM ba coded P2 adap e p ime we e
added in a o al eac ion olume o 10 μl and amplified acco ding o he ollowing he mal cycling p ofile:
72 °C o 3 min, 95 °C o 1 min, ollowed by 21 cycles o 95 °C o 10 s, 65 °C o 20 s and 72 °C o 3 min.
Pos amplifica ion he DNA concen a ion was de e mined using he Quan -I Picog een dsDNA assay
(The mo Fishe , Camb idge, UK). Lib a y DNA concen a ions ypically anged om 4 o 40 ng μl
1
(a e age o 16 ng μl
1
). Fo each sample om a 384 well pla e a 5 μl aliquo was pooled and spli in o wo
2 ml Lo bind Eppendo ubes (950 μl each). To each aliquo 950 μl o AMPu e XP (Beckman, High
Wycombe, UK) beads was added. Samples we e mixed, incuba ed (5 min, oom empe a u e) and placed
on a magne pa icle concen a o (MPC) un il he beads we e collec ed. The supe na an was disca ded.
The beads we e washed wice wi h 20 μl 70% e hanol while placed in he MPC and ai d ied (5 min). The
pooled lib a y was elu ed om he beads in 17 μl o 10 mM T is HCl pH 8.0. The wo 17 μl aliquo s o he
lib a y we e combined and he DNA concen a ion was de e mined using he Qbi de ice wi h he
Quan -I DNA HS Assay (In i ogen). Typical DNA concen a ions we e abo e 100 ng μl
1
. The DNA
size selec ion was pe o med using he Blue Pippin (Sage Science, Be e ly, MA, USA). Abou 3 μg o he
lib a y in 30 μl o 10 mM T is HCl pH 8.0 and 10 μl o he R2 ladde we e sepa a ed ( igh selec ion
p o ocol, 650 bp) using a 1.5% aga ose casse e acco ding o he manu ac u e ’s ins uc ions
(Sage Science, Be e ly, MA, USA), he eby yielding an a e age inse size o abou 485 bp. Size selec ed
samples we e collec ed in 40 μl o TRIS- TAPS bu e , pH 8.0 (Sage Science, Be e ly, MA, USA). The
a e age size o he lib a y was de e mined using a High Sensi i i y Chip and an Agilen 2100
Elec opho esis Bioanalyze (Agilen ). The DNA concen a ion was measu ed using he Qbi de ice and
he Quan -I DNA HS Assay (In i ogen). Size selec ed lib a ies we e quan ified using he Kappa
Biosciences Illumina lib a y qPCR quan ifica ion ki (Kapa Biosciences) on a S ep One qPCR machine
(The moFishe ) acco ding o he manu ac u e ’s ins uc ions and compa ed agains a known
concen a ion o a PhiX con ol lib a y. Se e al lib a ies we e pooled o sequencing in an equimola
manne , and he final pool was e-quan ified o sequencing ela i e o a s anda d lib a y o a known
concen a ion using he Kapa Biosciences Illumina lib a y qPCR quan ifica ion ki . Sequencing-
by-syn hesis o 6,144 BACs om ch omosome 2H was pe o med using an Illumina HiSeq2000 de ice
(2 × 100 cycles pai ed-end, single indexing ead, 384 BACs/lane) acco ding o manu ac u e ’s
ins uc ions, he eby yielding a leas 32 Gb/lane and an a e age sequence co e age o a leas 500- old
pe BAC. The emaining BAC clones om 2H (384 BACs/lane) and 0H (2304 BACs/lane) we e
sequenced wi h a HiSeq2500 machine (2 × 150 cycles pai ed-end, dual indexing, apid mode, yield: a
leas 30 Gb/lane) using a sligh ly adap ed p o ocol wi h an addi ional no maliza ion s ep p io o sample
pooling. B iefly, a cus om panel o 48 P5 and 48 P7 adap e oligos wi h 9 bp ba codes (wi h ≥4 hamming
dis ance) was designed o indi idually label up o 2,304 (48 × 48) lib a ies by dual indexing. A mix u e o
2μl o BAC DNA, 0.5 μl Nex e a 10 × Reac ion Bu e , 0.1 μl Nex e a Enzyme and 2.4 μl wa e was
incuba ed (5 min, 55 °C). Tn5 dena u a ion, eac ion clean-up, washing, elu ion and ans e o a esh
384 well pla e we e as desc ibed o he single-indexing lib a ies. 5 μl pu ified, agmen ed DNA, 2 μlo
5 × Kapa Robus 2G B Reac ion bu e , 0.2 μl o 10 mM dNTPs, 0.05 μl o Kapa Robus 2G Taq
polyme ase, 1 μl2μM P5 p ime , 1 μl2μM P7 p ime we e combined ( eac ion olume o 10 μl) and
amplified acco ding o ollowing he mal cycling p ofile: 72 °C o 3 min, 95 °C o 1 min, ollowed by
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 5
16 cycles o 95 °C o 10 s, 65 °C o 20 s and 72 °C o 3 min. The size p ofile and quan i y was
de e mined as desc ibed o single-indexing lib a ies. Amplified lib a ies we e no malised using
MagQuan bead echnology (GC Bio ech, Ne he lands) on a Calipe Zephy Robo (Pe kin Elme ),
essen ially as desc ibed by he manu ac u e . No malised lib a ies we e elu ed in 10 μl o 10 mM T is HCl
pH 8.0 and ans e ed o a esh 384 well PCR pla e.5 μl o 384 no malized samples we e pooled
( o al olume 1,920 μl). Pu ifica ion using AMPu e XP beads, washing, elu ion, size-selec ion
(Blue Pippin) and quali y checks p io o sequencing we e essen ially as desc ibed o single indexing
lib a ies. Sequencing-by-syn hesis o pooled lib a ies (2,304 BACs) was pe o med using an Illumina
HiSeq2500 de ice ( apid un mode, 2 × 150 cycles pai ed-end, dual indexing eads) acco ding o
manu ac u e ’s ins uc ions. A leas 40 Gbp/lane, and an a e age sequence co e age o >100- old pe
BAC we e ob ained (Da a Ci a ion 22, Da a Ci a ion 23, Da a Ci a ion 24, Da a Ci a ion 25).
Ma e-pai sequencing o MTP BACs
BAC clones we e inocula ed as desc ibed o he p epa a ion o sho gun lib a ies. The bac e ial cul u es
we e g own o 6 h a 37 °C in a shaking incuba o a 200 .p.m., and 384 clones we e pooled. The pool
was used o inocula e 250 ml 2 × YT media supplemen ed wi h chlo amphenicol (12.5 μgml
1
). The
cul u es we e incuba ed (18 h, 37 °C, 200 .p.m.). Cells we e ha es ed by cen i uga ion (3,220 g, 20 min,
4 °C), and he supe na an was disca ded. Alkali lysis and DNA isola ion s eps we e pe o med using he
La ge Cons uc ki (Qiagen, UK) essen ially ollowing he manu ac u e ’s ins uc ions. The DNA was
esuspended in 4.75 ml Bu e Ex, 100 μl 100 mM ATP (Fishe Scien ific, UK) we e added and
con amina ing E. coli DNA was emo ed using 150 μl ATP dependen Exonuclease (Qiagen). Du ing he
incuba ion (1 h, 37 °C) a Qiagen Tip-100 column (Qiagen) was equilib a ed in Bu e QBT (Qiagen). 5 ml
o Bu e QS we e added o he DNA, and he sample was applied o he equilib a ed column. The
column was washed wice wi h 10 ml o Bu e QC (Qiagen). The DNA was elu ed wi h 7.5 ml o
p e-wa med (65 °C) Bu e QF (Qiagen). The DNA was p ecipi a ed by adding 0.7 × olume o oom
empe a u e isop opanol and cen i uga ion (20 min, 3,220 g, 4 °C). The pelle was washed wice wi h
70% e hanol, ai d ied and dissol ed in 200 μl TE bu e acco ding o manu ac u e ’s guidelines. The
DNA concen a ion was measu ed using a Qubi Fluo ome e (The mo Fishe , Camb idge, UK) and
adjus ed wi h wa e o 13 ng μl
1
. Fo agmen a ion 200 μl dilu ed DNA we e equilib a ed (6 min, 55 °C)
and subsequen ly p o ided wi h 52 μl 5 × Tagmen Bu e Ma e-Pai and 8 μl Ma e-Pai Tagmen a ion
Enzyme (Illumina, San Diego, USA). A e he incuba ion (30 min, 55 °C), 65 μl Neu alize Tagmen
Bu e (Illumina, San Diego, USA) we e added, and he eac ion was incuba ed (5 min, oom
empe a u e). One olume CleanPCR beads (GC Bio ech, Alphen aan den Rijn, The Ne he lands) was
added, and he DNA was pu ified using magne ic sepa a ion. The DNA was elu ed in 170 μl o nuclease-
ee wa e , quan ified using a Qubi fluo ome e (DNA HS assay, In i ogen) and analysed using he
Agilen Bioanalyse (DNA 1,200 chip, Agilen , S ockpo , UK). S and displacemen was pe o med by
combining 105.3 μl o agmen ed DNA, 13 μl 10x S and Displacemen Bu e (Illumina), 5.2 μl dNTPs
(Illumina), 6.5 μl S and Displacemen Polyme ase (Illumina) and incuba ion (30 min, oom
empe a u e). CleanPCR beads (0.75 olume) we e added and he DNA was pu ified using a magne .
The DNA was elu ed in 30 μl nuclease- ee wa e . The concen a ion was measu ed (Qubi , DNA HS
assay, In i ogen), and a 1:6 dilu ed sample was analysed using he Agilen Bioanalyse (DNA 1,200 chip,
Agilen , S ockpo , UK). Size selec ion was pe o med using a Pippin Blue (Sage Science, Be e ly, MA,
USA). 30 μl DNA we e p o ided wi h 10 μl loading bu e and sepa a ed on a 0.75% aga ose casse e
(size selec ion cen e ed a 7 kb and collec ion be ween 6–8 kb) acco ding o he manu ac u e ’s
ins uc ions (Sage Science, Be e ly, MA, USA). Size selec ed samples we e collec ed in 40 μlo
TRIS- TAPS bu e (pH 8.0) (Sage Science, Be e ly, MA, USA), and analysed using he Agilen
Bioanalyse (high sensi i i y chip, Agilen , S ockpo , UK) o de e mine he final lib a y size. The DNA
concen a ion was measu ed using he Qubi de ice and he Quan -I DNA HS Assay (In i ogen).
Ci cula isa ion was pe o med by combining 40 μl size selec ed DNA, 12.5 μl 10 × ci cula isa ion bu e
(Illumina), 3 μl Ci cula isa ion Enzyme (Illumina) and 75 μl nuclease- ee wa e . The eac ion was
incuba ed a 30 °C o e nigh . Linea DNA was diges ed by adding 3.75 μl Exonuclease (Illumina) and
incuba ion (30 min, 37 °C). The enzyme was inac i a ed by hea (30 min, 70 °C) and he addi ion o 5 μl
s op liga ion (Illumina). Ci cula ised DNA (130 μl) was shea ed in a Co a is Mic oTube AFA Fibe
(P e-sli , Snap-cap, 6 × 16 mm; 2 cycles o 37 s, 10% du y cycle, 200 cycles pe bu s , 4 in ensi y, 4 °C)
using he Co a is S2 de ice (Co a is, Massachuse s, USA). M280 Dynabeads (The mo Fishe ) we e
p epa ed as desc ibed (Illumina). 130 μl washed M280 beads we e added o he agmen ed DNA, mixed
and placed on a lab o a o (20 min, oom empe a u e). Lib a y molecules we e a fini y pu ified and
washed as desc ibed (Illumina). The beads we e esuspended in a mix u e o 85 μl nuclease ee wa e ,
10 μl 10x End Repai Reac ion Bu e (Ilumina) and 5 μl end epai enzyme mix (Illumina) and incuba ed
(30 min, 30 °C). End epai ed lib a y molecules bound o M280 beads we e washed as desc ibed
(Illumina). A-Tailing and adap e liga ion we e pe o med acco ding o manu ac u e ’s ins uc ions
(Illumina). Fo PCR amplifica ion, he beads we e esuspended in a eac ion mix u e (20 μl nuclease- ee
wa e , 25 μl 2x Kappa HiFi (Kappa Biosys ems, London, UK), 5 μl Illumina P ime Cock ail) and
amplified (98 °C o 3 min, 12 cycles o 98 °C o 10 s, 60 °C o 30 s, 72 °C o 30 s ollowed by 72 °C o
5 min and s o age o he sample a 4 °C). Beads we e emo ed by magne ic sepa a ion and 45 μl o he
p oduc s we e ans e ed o a 2 ml DNA Lobind Eppendo ube. The DNA was p ecipi a ed by addi ion
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 6
o 31.5 μl CleanPCR beads (GC Bio ech, Alphen aan den Rijn, The Ne he lands). The beads we e washed
wice wi h 100 μl 70% e hanol, and he final lib a y was elu ed in 20 μl esuspension bu e (GC bio ech).
The DNA concen a ion was de e mined (Qubi , DNA HS assay, In i ogen), ollowed by analysis using
he Agilen Bioanalyse (High sensi i i y chip, Agilen , S ockpo , UK). Up o 12 ma e-pai lib a ies we e
pooled in an equimola manne and measu ed using he Kappa qPCR Illumina quan ifica ion ki .
Sequencing-by-syn hesis o pooled ma e-pai lib a ies was pe o med using an Illumina HiSeq2500
de ice ( apid un mode, 2 × 150 cycles pai ed-end, single indexing eads) acco ding o manu ac u e ’s
ins uc ions (Da a Ci a ion 26, Da a Ci a ion 27).
Sequence assembly o indi idual BACs
Assembly o gene-con aining BACs (UCR/JGI). A o al o 15,661 gene-bea ing BACs we e pai ed-end
sequenced (2 × 100 cycles) using he Illumina HiSeq2000 pla o m (Illumina, Inc., San Diego, CA, USA)
applying a combina o ial pooling design
24
, as desc ibed in Munoz-Ama iain e al.
13
. Reads we e
quali y immed, decon olu ed, and hen assembled BAC-by-BAC using Vel e e sion 1.2.09 ( e . 25)
wi h he pa ame e k se o 45. Sequences o an addi ional 50 andomly chosen BACs included in
Munoz-Ama iain e al.
13
we e de i ed using he Sange me hod by Jane G imwood (US Depa men o
Ene gy Join Genome Ins i u e) and Je emy Schmu z (HudsonAlpha Ins i u e o Bio echnology),
including sha e and ansposon sequencing. The assignmen o BACs o ch omosome a ms/pe i-
cen ome ic egions was pe o med using CLARK
26
, an accu a e k-me -based classifica ion me hod ha
is much as e han BLASTN o MegaBLAST. CLARK makes assignmen s by using a p ebuil da abase o
k-me s ha a e specific o each ch omosome a m/pe i-cen ome ic egion.
Assembly o MTP BACs om ba ley ch omosomes 1H, 3H, 4H, 6H and 7H (FLI and IPK). A o al
o 10,148 BACs mainly o igina ing om ba ley ch omosome 3H we e sequenced on he Roche 454
sys em. Reads we e decon olu ed and assigned o indi idual BACs
16
. Reads we e quali y immed
acco ding o he manu ac u e ’s ecommenda ions. Reads we e sc eened o E. coli and ec o sequences
wi h MegaBLAST
27
. Assemblies we e hen cons uc ed om he clean eads using he MIRA so wa e
28
as desc ibed in S eue nagel, e al.
16
and Taudien, e al.
29
.
A o al o 41,004 BACs we e sequenced on Illumina machines (mainly HiSeq2000) in pools o up o
672 indi idually ba coded BAC clones. Pai ed-end eads we e quali y immed wi h he CLC oolki and
sc eened o E. coli and ec o sequences wi h MegaBLAST. Assemblies we e ob ained by unning CLC
Assembly Cell Ve sion 4.0.6 be a wi h de aul pa ame e s. Con igs de i ed wi h low ead co e age as well
as con igs smalle han 500 bp we e emo ed using he c i e ia desc ibed in Beie , e al.
17
.
The esul an con igs we e hen compa ed o NCBI’s nucleo ide da abase using MegaBLAST o check
o possible con amina ion. Con igs wi h non-plan hi s we e ei he comple ely emo ed o immed.
Sca olding o MTP BACs om ba ley ch omosomes 1H, 3H, 4H, 6Hand7H(FLIandIPK).Sca olding
was pe o med as desc ibed in Beie e al.
17
B iefly, ma e-pai eads we e mapped agains he conca ena ed
assemblies o up o 384 BACs using BWA mem e sion 0.7.4 ( e . 30) wi h de aul pa ame e s. Only ead pai s
mapping uniquely (minimal mapping quali y o Q40) o di e en con igs o he same BAC assembly we e
e ained. These eads we e used o sca old indi idual BACs using SSPACE e sion 3.0 S anda d
31
.
I mul iple ma e-pai lib a ies we e p esen (MiSeq ma e-pai eads as well as HiSeq2000 ma e-pai
eads) an i e a i e sca olding p ocedu e
17
was used.
Assembly o MTP BACs om ba ley ch omosome 5H (BGI). Ob ained aw sequence eads om
5H MTP BACs we e fil e ed o gene a e high-quali y eads by he ollowing c i e ia: (1) eads con aining
mo e han 2% o Ns o wi h poly-A s uc u es we e emo ed; (2) eads wi h ≥40% low quali y bases o
sho inse size lib a ies (60% o la ge inse size lib a ies) we e excluded; (3) eads con aining adap e s
we e emo ed; (4) PCR duplica es we e de ec ed and excluded; (5) emo al o eads con amina ed by
E. coli, ec o sequences o phage sequences. High-quali y eads we e hen used o assembly.
BACs we e assembled using SOAPdeno o e sion 2.01 ( e . 32) mul iple imes using di e en k and m
alues (main pa ame e in SOAPdeno o assembly). In o al each BAC was assembled 45 imes (k om 33
o 66, only odd numbe s and m om 1 o 3). The N50 was examined o each assembly and he assembly
wi h he la ges N50 was e ained as he final assembly esul o each BAC.
Sca olding o MTP BACs om ba ley ch omosomes 5H (BGI). Assemblies om pai ed-end
sequences we e used as e e ence o mapping 2, 5 and 10 kb ma e-pai eads ob ained om ba ley
genomic WGS da a wi h SOAPaligne /soap2 e sion 2.21 wi h pa ame e s –p6– 3–R. Ma e-pai ead
pai s mapped in his ashion we e used in conjunc ion wi h he co esponding pai ed-end ead pai s o
e-assemble each BAC using SOAPdeno o e sion 2.01 as desc ibed abo e.
Assembly o MTP BACs om ba ley ch omosomes 2H and ‘0H’(EI). Minimal iling pa h BACs
om (i) ba ley ch omosomes 2H o om (ii) finge p in ed con igs no assigned o ch omosomes ( e med
‘0H’) we e sequenced. A e demul iplexing, sample quali y con ol (QC) in o ma ion was gene a ed
using Fas QC
33
. Con amina ion sc eening was ca ied ou using Kon aminan
34
. Reads we e sc eened
using a k-me size o 21 agains a ange o po en ial con aminan s (Phi X, E. coli,En e obac e cloacae
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 7
genomic DNA and BAC ec o ) and con amina ed eads o eads wi h quali y alues o30 we e
emo ed.
ABySS assemble ( 1.5.1)
35
was used o assemble he fil e ed pai ed-end eads o each BAC
indi idually (k-71, l-91 b-0). Pai ed-end con igs we e compa ed o NCBI’s NR da abase using BLAST o
check o hi s o non-plan o ganisms using e- alue 1e-4 as h eshold. The ob ained hi s we e compa ed
o NCBI axonomy using ‘ as acmd’ o ob ain common names used o check o any non-plan hi s.
Sca olding o MTP BACs om ba ley ch omosomes 2H and ‘0H’(EI). Illumina Nex e a ma e-pai
lib a ies we e c ea ed om pools o 384 BACs. A e quali y checking he eads using PAP
34
, he eads
we e me ged using FLASH ( e sion 1.2.9)
36
. Nex clip ( 0.8)
37
was un on he flashed eads o im he
junc ion adap e s. A k-me -based app oach was used o assign ma e-pai eads o indi idual BACs wi h
KAT ( 1.0.4) (h ps://gi hub.com/TGAC/KAT). Sca olding and gap closing we e pe o med on each
BAC indi idually using an in-house shell sc ip (a ailable om Gi Hub: h ps://gi hub.com/DhSaTGAC/
BAC-assembly-pipeline.gi ). SOAPdeno o sca olde e sion 2.01 ( e . 38) was applied o sca old he
ABySS pai ed-end con igs using he k-me classified ma e-pai eads wi h pa ame e s k=41, -G 30, -F, -w
and -L 100. The esul ing sca olds we e hen edi ed o eplace long s e ches (>20) o C/G wi h
‘N’cha ac e s as SOAP is known o subs i u e ‘N’s wi hin pai ed-end con igs o C/G. The sca olds we e
hen passed h ough GapClose ( 1.12- 6), a SOAP2 module, o fill in long s e ches o ‘N’s p oduced
du ing he sca olding s eps. Con igs and sca olds sho e han 500 bp we e emo ed o p oduce he final
assembly pe BAC.
Splash con amina ion checks o MTP BACs om ba ley ch omosomes 2H and ‘0H’(EI). The aw
eads wi hin each pla e we e aligned o one side o he ec o sequence adjacen o he es ic ion enzyme
cu si e using exone a e
39
. Subs ings o size 20 bp we e ex ac ed om aligning eads con aining he BAC
sequence adjacen o he ec o sequence. Flanking sequences om each BAC we e clus e ed based on a
Hamming dis anceo3 and consensus sequences gene a ed o accoun o sequencing e o s. These we e
compa ed wi h neighbo ing wells o check o po en ial con amina ion caused by splash du ing lab
p ocessing s eps. Whe e con amina ion be ween neighbo ing wells was indica ed, he assembled con igs
om each BAC in ques ion we e aligned in a pai wise ashion using exone a e and he o al pe cen age o
simila sequence (≥99% iden i y) was compu ed. In cases whe e neighbo ing BACs sha ed mo e han
10% simila sequence, bo h BACs we e esequenced.
Pseudomolecule cons uc ion
Ini ial con amina ion emo al. Sequence assemblies o 66,586 MTP clones, 5,468 non-MTP BACs
and 15,044 gene-bea ing clones
13
( o al numbe o unique BACs: 87,098) we e combined in o a single
FASTA file (Da a Ci a ion 28,Da a Ci a ion 29,Da a Ci a ion 30). I a clone had wo o mo e independen
sequence assemblies, we selec ed he one wi h he la ges N50 alue o u he analyses. BAC assemblies
we e aligned o a cus om lib a y o po en ial con aminan s (Da a Ci a ion 31) including phages, bac e ial
and ec o sequences using megablas
27
. Regions aligning o con aminan s (c i e ia: (alignmen
leng h ≥500 bp AND iden i y ≥80%) OR (iden i y ≥90%)) we e emo ed om he assembly using UNIX
sc ip s and BEDTools
40
. Sequences sho e han 500 bp o consis ing o less han 500 p ope nucleo ides
(ACGT cha ac e s) a e con amina ion emo al we e disca ded. This s ep emo ed 55.5 Mb (0.5%) o
he assembled BAC sequence.
Sequence alignmen o BACs sequences and o e lap de ec ion. A e con amina ion emo al, a se
o 87,075 BAC assemblies (Table 1, Da a Ci a ion 32) was aligned agains i sel using megablas
27
wi h a
wo d size o 44, e aining only alignmen s wi h iden i y ≥99% and alignmen leng h ≥500 bp. Two se s o
o e laps (s ingen and pe missi e) be ween BACs we e defined om he BLAST esul s o all BACs
agains each o he . Pai s o BACs we e conside ed as po en ially o e lapping unde s ingen c i e ia i
he e was a leas one high-sco ing pai (HSP) wi h alignmen leng h ≥5 kb and iden i y ≥99.8%. Unde
pe missi e c i e ia, we equi ed a leas one HSP wi h alignmen leng h ≥2 kb and iden i y ≥99.5%. Fo
all pai s o po en ially o e lapping BACs (unde ei he se o c i e ia), he size o hei o e lapping egions
was de e mined using UNIX sc ip s and BEDTools
40
as he ex en o non- edundan egions in he BAC
sequences (i.e., con igs o sca olds) con ained in HSPs ≥500 bp and iden i y ≥99.5% be ween BAC
sequences ha ing a leas one HSP wi h alignmen leng h ≥5 kb and iden i y ≥99.8% (s ingen c i e ia)
o alignmen leng h ≥2 kb and iden i y ≥99.5% (pe missi e c i e ia). HSPs less han 200 bp apa we e
combined in o one wi h BEDTools (command ‘me ge’). BAC o e lap in o ma ion was impo ed in o he
R s a is ical en i onmen
41
o use in gene ic ancho ing and me ging sequence assemblies o adjacen
BAC clones (see sec ion ‘Cons uc ion o he BAC o e lap g aph’).
Alignmen o BACs o he BioNano map o ba ley c . Mo ex. An op ical map o he genome o
ba ley c . Mo ex was gene a ed using he I ys pla o m o BioNano Genomics using N .BspQI as he
nicking enzyme. Fu he de ails o he op ical map p ocedu e a e desc ibed in Masche e al.
42
An in silico
BspQI diges was pe o med wi h he Knicke s so wa e (h p://www.bionanogenomics.com) using
de aul pa ame e s. Res ic ion maps o BAC sequences we e aligned o he BioNano map o ba ley c .
Mo ex
42
(Da a Ci a ion 33) wi h I ysView so wa e
43
(h p://www.bionanogenomics.com) using he
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 8
command line ool Re Aligne ( e sion 3827) wi h he ollowing pa ame e s ‘-M 2 -T 1e-4 -ex end 1
-biasw 0’ o epo all alignmen s wi h a confidence sco e ≥4.
Cons uc ion o he upda ed POPSEQ map o he Mo ex x Ba ke mapping popula ion. An ul a-
dense linkage map had been cons uc ed p e iously
5
by shallow whole-genome sho gun sequencing
o 90 ecombinan inb ed lines (RILs) de i ed om a c oss be ween he ba ley cul i a s Mo ex and Ba ke.
We wished o inc ease he esolu ion o his map by educing he a e age ac ion o missing da a pe
SNP ma ke . Towa ds his aim, we sequenced he exis ing Illumina pai ed-end lib a ies o 87 RILs
o highe co e age (2–3x) and combined hem (Da a Ci a ion 34) wi h he exis ing ead da a se
5
(ENA accession: ERP002184). Map cons uc ion ollowed he p ocedu es desc ibed in Chapman e al.
44
.
Reads we e aligned o he whole-genome sho gun assembly o ba ley c . Mo ex
4
(NCBI accession:
CAJW01) wi h BWA mem e sion 0.7.5a ( e . 45). So ing, con e sion o BAM o ma and emo al o
duplica e eads was done wi h Pica dTools e sion 1.100 (h p://b oadins i u e.gi hub.io/pica d/). Va ian
de ec ion and geno ype calling we e pe o med wi h SAMTools e sion 0.1.19 (commands ‘sam ools
mpileup –BD’and ‘bc ools iew –c g’). The esul an VCF file was fil e ed using an AWK sc ip
(Supplemen a y Tex S3 o Masche e al. 2013 ( e . 46)). Homozygous geno ype calls we e se o missing
i hei ead dep h was 0 o hei geno ype quali y below 3. He e ozygous geno ype calls we e se o
missing i hei ead dep h was below 3 o hei geno ype quali y below 5. Va ian s wi h (i) a quali y
sco es below 40, (ii) mo e han 10% he e ozygous geno ype calls, (iii) mo e han 90% missing da a a e
geno ype call fil e ing, o (i ) a mino allele equency below 5% we e disca ded. SNP in o ma ion was
agg ega ed a he con ig le el o de i e consensus geno ypes as desc ibed in he sec ion ‘F amewo k map
cons uc ion’in he Me hods sec ion o Chapman e al.
44
Fo map cons uc ion wi h MSTMap
47
, he
popula ion ype ‘RIL8’was used. Addi ional con igs we e inse ed in o he amewo k map as desc ibed
in Chapman e al.
44
(sec ion ‘Ancho ing sca olds on o he amewo k map’) using p e iously published
ead da a
5
. Va ian calling and map cons uc ion we e done o he O egon Wol e Ba ley (OWB) doubled
haploid popula ion using he same p ocedu es wi h he ollowing wo changes: (i) he e ozygous geno ype
calls we e excluded and (ii) he popula ion ype ‘DH’was used o map cons uc ion wi h MSTMap
47
.
Map posi ions in he OWB map we e in e pola ed in o he Mo ex x Ba ke map using loess eg ession in
R
41
. A consensus posi ion was de i ed as ollows: i map posi ions disag eed by mo e han 5 cM in bo h
maps, a con ig was conside ed unancho ed; o he wise, he Mo ex x Ba ke posi ion was p e e ed i
a ailable. The final map assigned gene ic posi ions o 791,176 WGS con igs (Table 2, Da a Ci a ion 35),
compa ed o 723,499 ancho ed con igs in he o iginal POPSEQ map
5
.
Gene ic ancho ing o single BAC clones. The gene ic posi ions o Mo ex WGS con igs in he upda ed
POPSEQ map we e li ed o BAC sequences ia sequence alignmen . The se o all con igs o he whole-
genome sho gun assembly o ba ley c . Mo ex
4
(NCBI accession: CAJW01) was aligned o all BAC
assemblies wi h megablas
27
using a wo d size o 44 and e aining only alignmen s wi h iden i y ≥99.8%
and alignmen leng h ≥1,000 bp. Fo each BAC clone, he gene ic posi ions o WGS con igs aligning o i s
cons i uen sequences we e abula ed and a gene ic posi ion o a clone was de i ed using a majo i y ule
wi h unc ions o he R package ‘da a. able’(h ps://c an. -p ojec .o g/web/packages/da a. able/index.
h ml). Nine y pe cen o con igs assigned o a BAC had o o igina e o he majo ch omosome and he
s anda d de ia ion o gene ic posi ions had o be ≤3 cM. BACs wi hou alignmen s o ancho ed WGS
con igs we e conside ed as unancho ed; hose no mee ing he consis ency c i e ia we e flagged as
‘inconsis en ly ancho ed’. In he second s ep, unancho ed clones we e posi ioned by u ilizing posi ional
in o ma ion om neighbo ing BACs. We conside ed as neighbo s o a gi en clone B all hose BACs ha
o e lapped o a leas 10% o hei assembled leng hs wi h clone B. The gene ic posi ion o an
MTP ch omosome no o . BACs in MTP no. o sequenced BACs no. o ancho ed BACs*a e age no. o sequences a e age N50 (kb)
1H 6,993 6,983 (99.9%) 6,410 (91.8%) 7.6 81.2
2H 9,061 8,969 (99.0%) 8,195 (91.4%) 9.9 104.5
3H 8,841 8,807 (99.6%) 8,303 (94.3%) 7.7 87.5
4H 8,314 8,306 (99.9%) 7,783 (93.7%) 6.7 91.2
5H 8,426 8,358 (99.2%) 7,573 (90.6%) 9.7 72.2
6H 8,305 7,886 (95.0%) 6,476 (82.1%) 7.4 70.7
7H 8,576 7,970 (92.9%) 6,842 (85.8%) 8.5 65.5
‘0H’
†
8,256 8,031 (97.3%) 6,714 (83.6%) 7.6 83.6
Non-MTP —21,765 20,397 (93.7%) 14.5 33.7
To al 66,772 87,075 78,693 (90.4%) 9.8 70.3
Table 1. BAC assembly and ancho ing s a is ics. *Numbe and pe cen age o BAC clones ha ha e been
assigned gene ic posi ions in he POPSEQ map.
†
BAC clones in physical con igs ha had no been assigned o
ch omosomes.
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 9
leng h ≥500 bp and an alignmen iden i y ≥99.5%. Regions o he old non- edundan sequence co e ed
by C (as de e mined by commands o BEDTools
40
sui e) we e emo ed and con ig C was added ins ead.
This p ocedu e was pe o med o each C2- ype con ig.
Nex , we que ied he GMAP alignmen s o genes ha had no alignmen s o he non- edundan
sequence, bu we e ep esen ed ei he in (a) he Mo ex WGS con igs o in (b) sequences o
BACs excluded om he o e lap analysis. We conside ed sequence o ype (a) and (b) as ‘addi ional
gene-bea ing sequences’. We aligned hese addi ional gene-bea ing sequences o he non- edundan
sequence wi h megablas
27
using a wo d size o 44 and conside ing only high-sco ing pai s wi h an
alignmen leng h ≥500 bp and an alignmen iden i y ≥99.5%. Regions co e ed by he non- edundan
sequence unde hese alignmen c i e ia we e sub ac ed om he addi ional gene-bea ing sequences and
sequence agmen s wi h a leng h ≥500 bp we e added o he non- edundan sequence.
Final con amina ion emo al. We iden ified egions in he non- edundan sequence ha we e no
co e ed by whole-genome sho gun eads o c . Mo ex. Alignmen o WGS eads and ead dep h
calcula ion we e done as desc ibed in he sec ion ‘Alignmen o Hi-C da a o es ic ion agmen s’.
Regions o he non- edundan sequence no co e ed by Mo ex WGS eads and wi h a leng h ≥500 bp
we e ex ac ed using UNIX command line ools and BEDTools
40
(command ‘ge as a’). The ex ac ed
sequences we e aligned o he NCBI NT da abase wi h megablas
27
using a wo d size o 44 and equi ing
he high-sco ing pai s o ha e a leng h o a leas 100 bp and an alignmen iden i y ≥80%. We e ained
only hi s whose desc ip ion in he NCBI NT da abase did no ma ch he ollowing egula exp ession
(R syn ax) ep esen ing a lis o common and axonomic names o plan species:
‘Ho deum|T i i|Populus|Aegilops|A ena|Alnus|A .squa osa|Mo us|Nelumbo|B assica|Cucumis|Ci us|
Camelina|F aga ia|Lo us|Ta enaya|Spa ina|Eucommia|So ghum|Co ylus|Theob oma|Phaseolus|Ba ley|
T i olium|Elymus|B achypodium|Be a ulga is|Ricinus|Licania|Phoenix|H . ulga e|Py us|Malus|P unus|
Saccha um|Hype icum|Whea |O yza|hlo oplas |Secale|Vi is|Que cus’
Regions o e lapping he BLAST hi s passing hese fil e s we e cu om he non- edundan sequence
wi h BEDTools
40
(command ‘sub ac ’). Sequences sho e han 500 bp a e he emo al o con aminan
sequences we e disca ded. This s ep emo ed 5 Mb (0.1%) o he assembled sequence.
Cons uc ion o pseudomolecule sequences o ch omosome 1H—7H and ch Un. We cons uc ed
pseudomolecules o he se en ba ley ch omosomes by placing he sequence agmen s o single BAC
assemblies ha cons i u e he non- edundan sequence acco ding o he Hi-C map posi ions o he BAC
o e lap clus e s hese agmen s belong o. Sequences no ancho ed by Hi-C we e placed on ch Un
(‘ch omosome unassigned’). The o de o clus e s was aken om he Hi-C map. BACs wi hin he same
clus e we e o de ed acco ding o he minimum spanning ee o he BAC o e lap g aph o he clus e
and o ien ed ela i e o he elome es using he Hi-C o ien a ion o he clus e i a ailable. The ela i e
o de o sequence agmen s o igina ing om he same BAC bin (see sec ion ‘Cons uc ion o he BAC
o e lap g aph’) could no be de e mined so ha he placemen o sequences wi hin a BAC bin (a e age
size: 70 kb) is a bi a y. Ch Un is composed o (i) sequence agmen s o igina ing om BAC o e lap
clus e s no placed in he Hi-C map, o (ii) gene-bea ing agmen s o BAC sequences and Mo ex WGS
con igs selec ed in addi ion o he non- edundan sequence (see sec ion Iden ifica ion o addi ional gene-
bea ing sequences). A gap o 100 N cha ac e s was inse ed be ween adjacen sequence agmen s.
Pseudomolecules o all ch omosomes and ch Un we e combined in o a single FASTA file (Da a Ci a ion
42). To accommoda e limi a ions o he Sequence/Alignmen Map o ma (see Usage No es) spli
pseudomolecules wi h a size below 512 Mb we e cons uc ed by b eaking pseudomolecules a bi a ily a
b eaks be ween sequence con igs (Da a Ci a ion 43, Da a Ci a ion 44). A BED file indica ing he
placemen o BAC sequence agmen s, Mo ex WGS con igs and in e cala ing gaps in he (spli )
pseudomolecules is a ailable o download (Da a Ci a ion 45, Da a Ci a ion 46).
A abula summa y o he posi ional in o ma ion inco po a ed in o pseudomolecules is gi en in Da a
Ci a ion 41.
Masking o esidual edundancy
Residual edundancy a ising om unde ec ed o e laps be ween adjacen BACs was de ec ed and masked
by aligning he pseudomolecules sequence o i sel wi h megablas
27
. Genomic in e als con ained in
BLAST hi s wi h a leng h ≥5 kb and an iden i y ≥99.8% we e conside ed as po en ially edundan (PR)
egions. PR egions we e classified o decide which sequence o a edundan pai o mask: (i) PR egions
assigned o ch omosomal pseudomolecules (as opposed o ch Un), bu ha ing BLAST hi s only o o he
ch omosomes we e conside ed as o igina ing om chime ic BAC assemblies inco po a ing un ela ed
sequences om di e en ch omosomes and masked wi h Ns; (ii) an analogous p ocedu es was used o
find in ach omosomal chime as based on Hi-C map in o ma ion; (iii) PR egions on ch Un ha had
alignmen s o egions on ch omosomal pseudomolecules we e masked, (i ) o o he PR egions one
sequence o a edundan pai was chosen a bi a ily. Posi ions o masked egions on he (spli )
pseudomolecules we e w i en in o a BED file (Da a Ci a ion 47, Da a Ci a ion 48). Masking was done
wi h BEDTools
40
(command ‘mask’) o e w i ing nucleo ides in edundan in e als wi h N cha ac e s.
Masked e sions o he (spli ) pseudomolecules a e p o ided as Da a Ci a ion 49, Da a Ci a ion 50).
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 16
POPSEQ gene ic map based on pseudomolecule sequence
A e he cons uc ion o he map-based e e ence sequence, we cons uc ed an upda ed high- esolu ion
gene ic map o he Mo ex x Ba ke popula ion o alida e he o de o gene ic map in he e e ence
Figu e 2. Collinea i y be ween he Hi-C map and wo gene ic maps. The posi ions o gene ic ma ke s
(x-axis) a e plo ed agains hei gene ic posi ions (y-axis) in a GBS map ( op ow) and a POPSEQ map
(bo om ow) o he Mo ex x Ba ke ecombinan inb ed lines.
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 17
sequence. Raw eads (see sec ion ‘Cons uc ion o he upda ed POPSEQ map o he Mo ex x Ba ke
mapping popula ion’) we e aligned o he ba ley pseudomolecules wi h BWA mem ( e sion 0.7.12)
45
.
Checking ma ed mapped pai ed eads, so ing, con e sion o BAM o ma and ma king o duplica e ead
pai s we e done wi h Pica dTools e sion 2.300 (h p://b oadins i u e.gi hub.io/pica d/). Va ian
de ec ion and geno ype calling we e pe o med using GATK Toolki e sion 3.3.0 (command
‘Haplo ypeCalle ’)
57
. A o al o fi e RILs wi h >3% he e ozygous a ian s we e emo ed. A a ian
posi ion was emo ed i mo e han 10% o all samples we e called he e ozygous, he e we e mo e han
80% missing da a, o he mino allele equency (in he non-missing da a) was smalle han 5%. SNP
in o ma ion was agg ega ed a he con ig le el o de i e consensus geno ype blocks wi h alse disco e y
a e calcula ed based on he quali y o each a ian call in he block. High-confidence geno ype blocks
we e ob ained based on a Bon e oni co ec ion h eshold. Gi en he ac ha he leng h o c osso e
ac s is significan ly la ge han ha o non-c osso e ac s and non-c osso e ac s would enla ge he
gene ic dis ance a ificially, we only e ained high-confidence geno ype blocks wi h mo e han 1 Mb ac
leng h, which a e likely o be de i ed om c osso e s. Rep esen a i e non- edundan genomic a ian s o
high-confidence geno ype blocks we e ex ac ed and used o he cons uc ion o a high- esolu ion map
h ough MSTMap
47
. We u he ancho ed all emaining ma ke s o he gene ic map by he C p og am
‘cancho ’
5
. The final POPSEQ map consis ed o 9,012,742 SNP a ian s defined on he pseudomolecule
sequence Da a ci a ion 51).
Rep esen a ion o ull-leng h cDNAs
The ep esen a ion o gene models in he whole-genome genome assembly o ba ley c . Mo ex
4
and in
he pseudomolecules was compa ed by aligning a se o 22,651 publicly a ailable ull-leng h cDNAs
55
o
he assemblies using he GMAP splice aligne so wa e
56
. The GMAP alignmen ou pu was hen fil e ed.
I a ull-leng h cDNA had mul iple hi s, only he hi wi h he highes % iden i y was conside ed. Hi s we e
u he fil e ed by iden i y (≥98%) and co e age ( ≥95%). This esul ed in a se o hi s ep esen ing
genes eco e ed in ac on a single genomic con ig/ch omosome.
Code a ailabili y
R and shell sou ce code o he cons uc ion o he BAC o e lap g aph and he Hi-C map is p o ided as
Da a Ci a ion 52. Code can be e-used unde he e ms o he MIT license.
Da a Reco ds
BAC sequence aw da a was submi ed o he Eu opean Nucleo ide A chi e (ENA) (Da a Ci a ion 1,
Da a Ci a ion 2, Da a Ci a ion 3, Da a Ci a ion 4, Da a Ci a ion 5, Da a Ci a ion 6, Da a Ci a ion 7,
Da a Ci a ion 8, Da a Ci a ion 9, Da a Ci a ion 10, Da a Ci a ion 11, Da a Ci a ion 12, Da a Ci a ion 13,
Da a Ci a ion 14, Da a Ci a ion 15, Da a Ci a ion 16, Da a Ci a ion 17, Da a Ci a ion 18, Da a Ci a ion 19,
Da a Ci a ion 20, Da a Ci a ion 21, Da a Ci a ion 22, Da a Ci a ion 23, Da a Ci a ion 24,
Da a Ci a ion 25, Da a Ci a ion 26, Da a Ci a ion 27). BAC assemblies we e submi ed o ENA o NCBI
(Da a Ci a ion 28, Da a Ci a ion 29). Raw da a o POPSEQ (Da a Ci a ion 35), GBS (Da a Ci a ion 38) and
Hi-C mapping (Da a Ci a ion 40) we e submi ed o ENA. P ocessed da ase s a e accessible as
Figu e 3. Collinea i y be ween he Hi-C map and a cy ogene ic map o ch omosome 3H. Do s ma k he
posi ions o p obes in he cy ogene ic map (x-axis) and he Hi-C-de i ed pseudomolecule (y-axis). A linea
eg ession line ( ed) was fi ed wi h he R unc ion lm(). No e ha cy ogene ic da a is no a ailable o dis al
egions because p obes we e designed only o non- ecombining pe i-cen ome ic egions
61
.
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 18
QRLPDAAGSSAEEHSGQDKLLIVVTPTR
ARASQAYYLSRMGQTLRLVRPPVLWVVV
EAGKPTPEAALELRRTAVMHRYVGCCDA
LNASASPAVDFRPHQLNAGLEVVENHRL
DGVVYFADEEGVYSLPLFDRLRQIRRFG
TWPVPTISDGGHGVVLEGPVCKQNQVVG
WHTSGDANKLQRFHVAMSGFAFNSTMLW
DPRLRSHKAWNSIRHPEMVEQGFQGTTF
VEQLVEDESQMEGIPADCSQIMNWHVPF
GSESPVYPKGWRSAANLDVIIPLK
Figu e 4. Accessing sequence and posi ional in o ma ion wi h he ba ley genome explo e (BARLEX). The
ba ley pseudomolecule da a was impo ed in o BARLEX, whe e i is di ec ly linked o he IPK Ba ley BLAST
se e . Use s can pas e a nucleo ide o amino acid sequence (1) in o he BARLEX inpu que y o m and selec
e e ence da abase such as pseudomolecules sequence, he se o all BAC assemblies o anno a ed genes (2). The
sequence is hen ans e ed o he IPK ba ley BLAST Se e (3). The web page wi h he BLAST esul s (4)
con ains e e ences o BARLEX in o ma ion pages o di e en s uc u al uni s (BAC sequence con igs, BAC,
BAC clus e , ch omosomal Hi-C map). Fo example, he pages o BAC sequence con igs isualize he epea
con en based on genome-wide k-me his og ams (5) and a e linked o a g aph-based isualiza ion (6) o he
en i e BAC assembly. Summa y s a is ics and posi ional in o ma ion o BAC clus e s a e p esen ed in ables
ha can be sea ched, so ed and subse ed using use -defined c i e ia (7). Use s can con e pseudomolecule
coo dina es (AGP posi ions) o in e als in he unde lying BAC sequence assemblies (8).
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 19
Digi al Objec Iden ifie s (DOIs) in he Plan Genomics and Phenomics Resea ch Da a Reposi o y
58
(Da a
Ci a ion 30, Da a Ci a ion 31, Da a Ci a ion 32, Da a Ci a ion 33, Da a Ci a ion 34, Da a Ci a ion 36, Da a
Ci a ion 37, Da a Ci a ion 39, Da a Ci a ion 41, Da a Ci a ion 42, Da a Ci a ion 43, Da a Ci a ion 44, Da a
Ci a ion 45, Da a Ci a ion 46, Da a Ci a ion 47, Da a Ci a ion 48, Da a Ci a ion 49, Da a Ci a ion 50, Da a
ci a ion 51, Da a Ci a ion 52). DOIs we e egis e ed wi h e!DAL
59
.
Technical Valida ion
Collinea i y be ween gene ic maps and pseudomolecules
To alida e he o de o sca olds in he Hi-C map, we compa ed he o de o gene ic ma ke loci in he
Hi-C-de i ed pseudomolecules o hei posi ions in linkage maps. Fi s , we used geno yping-by-
sequencing (GBS)
11,50
o ype single-nucleo ide polymo phisms (SNPs) seg ega ing in a bi-pa en al
popula ion comp ising 2,398 ecombinan inb ed lines (RILs). A o al o 2,637 SNPs we e de ec ed by
aligning GBS eads and calling a ian s and geno ypes using a p e iously published pipeline
46
. Second, we
eanalysed WGS e-sequencing da a o a subse o he same popula ion (POPSEQ da a) comp ising 90
RILs. Cons uc ion o a amewo k linkage map and inse ion o addi ional ma ke s we e pe o med
essen ially as desc ibed by Chapman e al.
44
A do plo compa ison o physical and gene ic SNP posi ions
e ealed ha ma ke o de s we e highly collinea be ween he pseudomolecules and bo h he GBS and
POPSEQ map o he Mo ex x Ba ke popula ion (Fig. 2).
Collinea i y be ween a cy ogene ic map and he pseudomolecule o ch omosome 3H
We could no alida e he o de o BAC o e lap clus e s in he la ge pe i-cen ome ic egions because o
se e ely ep essed ecombina ion
3,60
. The e o e, we compa ed he o de o p obes mapped by
fluo escence in-si u hyb idiza ion o ch omosomal loca ions on ch omosome 3H and hei co esponding
sequences in he pseudomolecule o 3H. Since p obes we e de i ed om BAC sequences associa ed wi h
physical con igs, hei posi ion om he e e ence sequence could be de e mined om he BAC o e lap
g aph. The compa ison showed ha he cy ogene ic and Hi-C maps we e highly collinea in pe i-
cen ome ic egions o ch omosome 3H (Fig. 3).
Rep esen a ion o ull-leng h cDNAs
To assess he comple eness o ou assembly, we checked o he p esence o high-confidence ansc ip
sequences. The ep esen a ion o gene models in he whole-genome sho gun assembly o ba ley c .
Mo ex
4
and in he map-based e e ence assembly was compa ed by aligning a se o 22,651 publicly
a ailable ull-leng h cDNAs
55
o ba ley c . ‘Ha una Nijo’. A e aligning and fil e ing, 18,062 (79.74%)
in ac ull-leng h cDNAs we e ound in he pseudomolecules, whe eas only 10,496 (46.33%) we e
eco e ed in he whole-genome assembly. This inc ease in he numbe o co ec ly ep esen ed ull-leng h
cDNAs indica es he e o in es ed in he map-based assembly. Ne e heless, a significan p opo ion o
genes emain agmen ed e en in he pseudomolecule assembly (20.26%), and p esumably hese la gely
ep esen di ficul o assemble genes ha con ain e.g., mic osa elli es, long homopolyme s e ches and
o he di ficul ea u es, and/o o m pa o complex gene amilies ha a e di ficul o esol e. I is likely
ha only longe ead echnologies such as Pacific Biosciences (h p://www.pacb.com) o Ox o d
Nanopo e (h ps://www.nanopo e ech.com) will be able o esol e hese mo e di ficul cases. Fu he
esul s on gene space comple eness based on an au oma ed gene anno a ion o he pseudomolecules, and
on he ep esen a ion o epe i i e elemen s a e desc ibed elsewhe e
42
.
Usage No es
Posi ional in o ma ion o BAC sequences, physical con igs and WGS con igs can be accessed ia he
ba ley genome explo e BARLEX (Fig. 4). BLAST sea ches agains he ba ley pseudomolecules can also be
ca ied ou in BARLEX. We no e ha p ocessing BAM files wi h sho ead alignmen s o he ull
pseudomolecules wi h commonly used ools such as SAM ools
52
o BEDTools
40
may no wo k as
expec ed because o es ic ions on he ch omosome size (512 Mb) o indexing file in Sequence
Alignmen /Map (SAM) o ma
52
. To ci cum en his issue, we ha e spli he pseudomolecules in o wo
pa and p o ide (i) a FASTA file wi h spli pseudomolecules (Da a Ci a ion 44) along he wi h he in ac
sequences and (ii) a BEDfile o con e be ween ull and spli pseudomolecule coo dina e (Da a Ci a ion
43) Al e na i ely, he CRAM o ma (h ps://sam ools.gi hub.io/h s-specs/CRAM 3.pd ) may be used
ins ead o he BAM o ma . We no e ha he o ien a ion o sequence con igs wi hin indi idual BACs in
he pseudomolecules is a bi a y, hus he o de and o ien a ion o sequences in he pseudomolecules is
accu a e only up o esolu ion o ~100 kb.
Re e ences
1. Schul e, D. e al. The in e na ional ba ley sequencing conso ium--a he h eshold o e ficien access o he ba ley genome. Plan
physiology 149, 142–147 (2009).
2. Schul e, D. e al. BAC lib a y esou ces o map-based cloning and physical map cons uc ion in ba ley (Ho deum ulga e L).
BMC genomics 12, 247 (2011).
3. A iyadasa, R. e al. A sequence- eady physical map o ba ley ancho ed gene ically by wo million single-nucleo ide poly-
mo phisms. Plan physiology 164, 412–423 (2014).
4. In e na ional Ba ley Genome Sequencing Conso ium. A physical, gene ic and unc ional sequence assembly o he
ba ley genome. Na u e 491, 711–716 (2012).
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 20
5. Masche , M. e al. Ancho ing and o de ing NGS con ig assemblies by popula ion sequencing (POPSEQ). The Plan Jou nal 76,
718–727 (2013).
6. Lande , E. S. e al. Ini ial sequencing and analysis o he human genome. Na u e 409, 860–921 (2001).
7. Schnable, P. S. e al. The B73 maize genome: complexi y, di e si y, and dynamics. Science 326, 1112–1115 (2009).
8. Lam, E. T. e al. Genome mapping on nanochannel a ays o s uc u al a ia ion analysis and sequence assembly. Na u e
bio echnology 30, 771–776 (2012).
9. Liebe man-Aiden, E. e al. Comp ehensi e mapping o long- ange in e ac ions e eals olding p inciples o he human genome.
Science 326, 289–293 (2009).
10. Bu on, J. N. e al. Ch omosome-scale sca olding o de no o genome assemblies based on ch oma in in e ac ions. Na u e
bio echnology 31, 1119–1125 (2013).
11. Poland, J. A., B own, P. J., So ells, M. E. & Jannink, J.-L. De elopmen o high-densi y gene ic maps o ba ley and whea using
a no el wo-enzyme geno yping-by-sequencing app oach. PLoS ONE 7, e32253 (2012).
12. Colmsee, C. e al. BARLEX— he Ba ley D a Genome Explo e . Mol Plan 8, 964–966 (2015).
13. Munoz-Ama iain, M. e al. Sequencing o 15 622 gene-bea ing BACs cla ifies he gene-dense egions o he ba ley genome. Plan
Jou nal 84, 216–227 (2015).
14. Pasqua iello, M. e al. The ba ley F os esis ance-H2 locus. Func ional & in eg a i e genomics 14, 85–100 (2014).
15. Meye , M., S enzel, U. & Ho ei e , M. Pa allel agged sequencing on he 454 pla o m. Na u e p o ocols 3, 267–278 (2008).
16. S eue nagel, B. e al. De no o 454 sequencing o ba coded BAC pools o comp ehensi e gene su ey and genome analysis in he
complex genome o ba ley. BMC genomics 10, 547 (2009).
17. Beie , S. e al. Mul iplex sequencing o bac e ial a ificial ch omosomes o assembling complex plan genomes. Plan bio-
echnology jou nal 14, 1511–1522 (2016).
18. Samb ook, J. & Russell, D. W. Molecula cloning: a labo a o y manual. 3 d edi ion (Coldsp ing-Ha bou Labo a o y P ess, 2001).
19. Ai d, D. e al. Analyzing and minimizing PCR amplifica ion bias in Illumina sequencing lib a ies. Genome biology 12, R18
(2011).
20. Quail, M. A. e al. A la ge genome cen e ’s imp o emen s o he Illumina sequencing sys em. Na u e me hods 5,
1005–1010 (2008).
21. Asan e al. Pai ed-end sequencing o long- ange DNA agmen s o de no o assembly o la ge, complex Mammalian genomes by
di ec in a-molecule liga ion. PLoS ONE 7, e46211 (2012).
22. Meye , M. & Ki che , M. Illumina sequencing lib a y p epa a ion o highly mul iplexed a ge cap u e and sequencing. Cold
Sp ing Ha b P o oc 2010, pdb p o 5448 (2010).
23. Adey, A. e al. Rapid, low-inpu , low-bias cons uc ion o sho gun agmen lib a ies by high-densi y in i o ansposi ion.
Genome biology 11, R119 (2010).
24. Lona di, S. e al. Combina o ial pooling enables selec i e sequencing o he ba ley gene space. PLoS compu a ional biology 9,
e1003010 (2013).
25. Ze bino, D. R. & Bi ney, E. Vel e : algo i hms o de no o sho ead assembly using de B uijn g aphs. Genome esea ch 18,
821–829 (2008).
26. Ouni , R., Wanamake , S., Close, T. J. & Lona di, S. CLARK: as and accu a e classifica ion o me agenomic and genomic
sequences using disc imina i e k-me s. BMC genomics 16, 236 (2015).
27. Zhang, Z., Schwa z, S., Wagne , L. & Mille , W. A g eedy algo i hm o aligning DNA sequences. Jou nal o compu a ional
biology: a jou nal o compu a ional molecula cell biology 7, 203–214 (2000).
28. Che eux, B., We e , T. & Suhai, S. in Ge man con e ence on bioin o ma ics (1999); 45–56.
29. Taudien, S. e al. Sequencing o BAC pools by di e en nex gene a ion sequencing pla o ms and s a egies. BMC esea ch no es
4, 411 (2011).
30. Li, H. & Du bin, R. Fas and accu a e sho ead alignmen wi h Bu ows–Wheele ans o m. Bioin o ma ics 25,
1754–1760 (2009).
31. Boe ze , M., Henkel, C. V., Jansen, H. J., Bu le , D. & Pi o ano, W. Sca olding p e-assembled con igs using SSPACE.
Bioin o ma ics 27, 578–579 (2011).
32. B enchley, R. e al. Analysis o he b ead whea genome using whole-genome sho gun sequencing. Na u e 491, 705–710 (2012).
33. And ews, S. Fas QC: a quali y con ol ool o high h oughpu sequence da a. A ailable online a : h p://www.bioin o ma ics.
bab aham.ac.uk/p ojec s/ as qc (2010).
34. Legge , R. M., Rami ez-Gonzalez, R. H., Cla ijo, B. J., Wai e, D. & Da ey, R. P. Sequencing quali y assessmen ools o enable
da a-d i en in o ma ics o high h oughpu genomics. F on ie s in gene ics 4, 288 (2013).
35. Simpson, J. T. e al. ABySS: a pa allel assemble o sho ead sequence da a. Genome esea ch 19, 1117–1123 (2009).
36. Magoc, T. & Salzbe g, S. L. FLASH: as leng h adjus men o sho eads o imp o e genome assemblies. Bioin o ma ics 27,
2957–2963 (2011).
37. Legge , R. M., Cla ijo, B. J., Clissold, L., Cla k, M. D. & Caccamo, M. Nex Clip: an analysis and ead p epa a ion ool o Nex e a
Long Ma e Pai lib a ies. Bioin o ma ics 30, 566–568 (2014).
38. Luo, R. e al. SOAPdeno o2: an empi ically imp o ed memo y-e ficien sho - ead de no o assemble . GigaScience 1, 18 (2012).
39. Sla e , G. S. & Bi ney, E. Au oma ed gene a ion o heu is ics o biological sequence compa ison. BMC bioin o ma ics 6,
31 (2005).
40. Quinlan, A. R. & Hall, I. M. BEDTools: a flexible sui e o u ili ies o compa ing genomic ea u es. Bioin o ma ics 26,
841–842 (2010).
41. R: A Language and En i onmen o S a is ical Compu ing (R Founda ion o S a is ical Compu ing, 2015).
42. Masche , M. e al. A ch omosome con o ma ion cap u e o de ed sequence o he ba ley genome. Na u e doi:10.1038/na u e22043
(2017).
43. Cao, H. e al. Rapid de ec ion o s uc u al a ia ion in a human genome using nanochannel-based genome mapping echnology.
GigaScience 3, 1 (2014).
44. Chapman, J. A. e al. A whole-genome sho gun app oach o assembling and ancho ing he hexaploid b ead whea genome.
Genome biology 16, 26 (2015).
45. Li, H. Aligning sequence eads, clone sequences and assembly con igs wi h BWA-MEM. P ep in a h ps://a xi .o g/pd /
1303.3997 2.pd (2013).
46. Masche , M., Wu, S., Amand, P. S., S ein, N. & Poland, J. Applica ion o geno yping-by-sequencing on semiconduc o sequencing
pla o ms: a compa ison o gene ic and e e ence-based ma ke o de ing in ba ley. PLoS ONE 8, e76925 (2013).
47. Wu, Y., Bha , P. R., Close, T. J. & Lona di, S. E ficien and accu a e cons uc ion o gene ic linkage maps om he minimum
spanning ee o a g aph. PLoS gene ics 4, e1000212 (2008).
48. Csa di, G. & Nepusz, T. The ig aph so wa e package o complex ne wo k esea ch, In e Jou nal, Complex Sys ems 1695 (2006).
49. P im, R. C. Sho es connec ion ne wo ks and some gene aliza ions. Bell sys em echnical jou nal 36, 1389–1401 (1957).
50. Wendle , N. e al. Unlocking he seconda y gene-pool o ba ley wi h nex -gene a ion sequencing. Plan bio echnology jou nal 12,
1122–1131 (2014).
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 21
51. Ma in, M. Cu adap emo es adap e sequences om high- h oughpu sequencing eads. EMBne . jou nal 17, 10–12 (2011).
52. Li, H. e al. The Sequence Alignmen /Map o ma and SAM ools. Bioin o ma ics 25, 2078–2079 (2009).
53. Ga ison, E. & Ma h, G. Haplo ype-based a ian de ec ion om sho - ead sequencing. P ep in a h ps://a xi .o g/pd /
1207.3907 2.pd (2012).
54. Kalho , R., Tjong, H., Jaya hilaka, N., Albe , F. & Chen, L. Genome a chi ec u es e ealed by e he ed ch omosome con o ma ion
cap u e and popula ion-based modeling. Na u e bio echnology 30, 90–98 (2012).
55. Ma sumo o, T. e al. Comp ehensi e sequence analysis o 24,783 ba ley ull-leng h cDNAs de i ed om 12 clone lib a ies. Plan
physiology 156, 20–28 (2011).
56. Wu, T. D. & Wa anabe, C. K. GMAP: a genomic mapping and alignmen p og am o mRNA and EST sequences. Bioin o ma ics
21, 1859–1875 (2005).
57. DeP is o, M. A. e al. A amewo k o a ia ion disco e y and geno yping using nex -gene a ion DNA sequencing da a. Na u e
gene ics 43, 491–498 (2011).
58. A end, D. e al. PGP eposi o y: a plan phenomics and genomics da a publica ion in as uc u e. Da abase 2016, baw033 (2016).
59. A end, D. e al. e!DAL--a amewo k o s o e, sha e and publish esea ch da a. BMC bioin o ma ics 15, 214 (2014).
60. Künzel, G., Ko zun, L. & Meis e , A. Cy ologically in eg a ed physical es ic ion agmen leng h polymo phism maps o he
ba ley genome based on ansloca ion b eakpoin s. Gene ics 154, 397–412 (2000).
61. Aliye a-Schno , L. e al. Cy ogene ic mapping wi h cen ome ic bac e ial a ificial ch omosomes con igs shows ha his
ecombina ion-poo egion comp ises mo e han hal o ba ley ch omosome 3H. The Plan Jou nal 84, 385–394 (2015).
Da a Ci a ions
1. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9062 (2016).
2. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9097 (2016).
3. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9098 (2016).
4. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9099 (2016).
5. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9100 (2016).
6. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9101 (2016).
7. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9102 (2016).
8. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9103 (2016).
9. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9104 (2016).
10. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8576 (2016).
11. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8577 (2016).
12. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8578 (2016).
13. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9619 (2016).
14. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8579 (2016).
15. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB8580 (2016).
16. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9429 (2016).
17. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9430 (2016).
18. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9431 (2016).
19. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB10963 (2016).
20. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11489 (2016).
21. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB12096 (2016).
22. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11758 (2016).
23. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9428 (2016).
24. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11991 (2016).
25. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB9427 (2016).
26. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11798 (2016).
27. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB11992 (2016).
28. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB13020 (2016).
29. Muñoz-Ama iaín, M. e al. NCBI BioP ojec PRJNA198204 (2015).
30. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/21 (2016).
31. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/28 (2016).
32. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/12 (2016).
33. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/31 (2016).
34. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB13028 (2016).
35. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/33 (2016).
36. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/22 (2016).
37. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/30 (2016).
38. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB14130 (2016).
39. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/29 (2016).
40. In e na ional Ba ley Genome Sequencing Conso ium. Eu opean Nucleo ide A chi e PRJEB14169 (2016).
41. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/20 (2016).
42. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/34 (2016).
43. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/27 (2016).
44. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/36 (2016).
45. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/23 (2016).
46. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/24 (2016).
47. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/25 (2016).
48. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/26 (2016).
49. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/35 (2016).
50. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/37 (2016).
51. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/17 (2016).
52. In e na ional Ba ley Genome Sequencing Conso ium. IPK Ga e sleben h p://dx.doi.o g/10.5447/IPK/2016/19 (2016).
Acknowledgemen s
This wo k was ca ied ou unde he auspices o he In e na ional Ba ley Genome Sequencing
Conso ium and suppo ed om he ollowing unding sou ces: Ge man Minis y o Educa ion and
Resea ch (BMBF) g an 0314000 ‘BARLEX’and 0315954 ‘TRITEX’ o M.P., U.S. and N.S and 031A536
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 22
‘de.NBI’ o U.S. Leibniz Associa ion g an (‘Pak . Fo schung und Inno a ion’)‘sequencing ba ley
ch omosome 3H’ o N.S. and U.S.; Sco ish Go e nmen /UK Bio echnology and Biological Sciences
Resea ch Council (BBSRC) g an BB/100663X/1 o R.W, P.E.H., J.R.; BBSRC g an s BB/I008357/1 o M.
D.C., M.C. and BB/I008071/1 o P.K.; o Finland g an 266430 and a BioNano g an o A.H.S.; Ca lsbe g
Founda ion g an n . 2012_01_0461 o he Ca lsbe g Resea ch Labo a o y; G ain Resea ch and
De elopmen Co po a ion (GRDC) g an DAW00233 o C.L. and P.L.; Depa men o Ag icul u al and
Food, Go e nmen o Wes e n Aus alia g an 681 o C.L.; Na ional Na u al Science Founda ion o China
(NSFC) g an 31129005 o C.L. and G.Zhang; NSFC g an 31330055 o G.Zhang.; Czech Minis y o
Educa ion, You h and Spo s g an LO1204 o J.D.; Na ional Science Founda ion g an DBI 0321756
‘Coupling EST and Bac e ial A ificial Ch omosome Resou ces o Access he Ba ley Genome’ o T.J.C.
and S.L.; Uni ed S a es Depa men o Ag icul u e (USDA), Ag icul u e and Food Resea ch Ini ia i e
Plan Genome, Gene ics and B eeding P og am o USDA-CSREES-NIFA g an 2009-65300-05645
‘Ad ancing he Ba ley Genome’and 2011-68002-30029 ‘T i iceaeCAP’ o T.J.C., S.L. and G.J.M.; Uni ed
S a es Na ional Science Founda ion (NSF)-ABI g an DBI-1062301 o T.J.C. and S.L.; Uni e si y o
Cali o nia g an CA-R-BPS-5306-H o T.J.C and S.L.;Na ional Science Founda ion g an DBI 0321756
‘Algo i hms o Genome Assembly o Ul a-deep Sequencing Da a’ o S.L. Nex -gene a ion sequencing
and lib a y cons uc ion was deli e ed ia he BBSRC Na ional Capabili y in Genomics
(BB/J010375/1) a Ea lham Ins i u e ( o me ly The Genome Analysis Cen e) by membe s o he
Pla o ms and Pipelines g oup and BBSRC Ins i u e S a egic P og amme unding o Bioin o ma ics
(BB/J004669/1) o M.D.C., S.A. and M.C. We g a e ully acknowledge: (1) he excellen echnical
assis ance by Susanne König, Manuela Knau , Uli Beie , Anne Kusse ow, Ka in T nka, Ines Walde,
Sand a D iesslein, Cyn hia Voss; (2) Do een S engel, Anne Fiebig, Thomas Münch, Danu a Schüle
and Daniel A end and Ma hias Lange o sequence aw da a managemen and da a submission o
EMBL/ENA and egis a ion o DOIs; (3) D Hélène Be ges, A naud Bellec and Sonia Vau in (CNRGV)
o managemen and dis ibu ion o ba ley BAC lib a ies; (4) And eas G ane and Da id Ma shall o
scien ific discussions.
Au ho Con ibu ions
BAC sequencing and assembly (1H, 3H, 4H): S.B., A.Himmelbach, S.T., M.F., M.G., M.M., U.S.
(co-leade ), M.P. (co-leade ), N.S. (leade ); BAC sequencing and assembly (2H, unassigned): D.S., D.H., S.
A. (co-leade ), M.D.C. (co-leade ), M.C. (co-leade ), R.W. (leade ); BAC sequencing and assembly
(5H, 7H): X.Z., R.A.B., Q.Z., C.T., J.K.M., B.C., G.Zhou, F.D., Y.H., S.Y., S.Cao, S.Wang, X.L., M.I.B., P.L.,
G.Zhang (co-leade ), C.Li (leade ); BAC sequencing and assembly (6H): S.B., S.Wang, C.Lin, H.L., U.S., M.
H. (co-leade ), I.B. (leade ); BAC sequencing (gene-bea ing): M.M.-A., R.O., S.Wanamake , S.L.
(co-leade ), T.J.C. (leade ); Op ical mapping: A.Has ie, H.Š., J.T., H.S., J.V., S.Chan, M.M., N.S., J.D.,
A.H.S. (leade ); Ch omosome con o ma ion cap u e: A.Himmelbach, S.G., M.M. (co-leade ), N.S. (leade );
Pseudomolecule cons uc ion: M.M. (leade ), S.B., C.C., D.B., T.S., P.K., N.S., U.S. (co-leade ); Valida ion:
L.L., M.B., L.A.-S., A.Houben, J.A.P., N.S., G.J.M., M.M. (leade ). All au ho s ead and commen ed on he
manusc ip .
Addi ional in o ma ion
Compe ing financial in e es s: The au ho s decla e no compe ing financial in e es s.
How o ci e his a icle: Beie , S. e al. Cons uc ion o a map-based e e ence genome sequence o
ba ley, Ho deum ulga e L. Sci. Da a 4:170044 doi: 10.1038/sda a.2017.44 (2017).
Publishe ’s no e: Sp inge Na u e emains neu al wi h ega d o ju isdic ional claims in published maps
and ins i u ional a filia ions.
This wo k is licensed unde a C ea i e Commons A ibu ion 4.0 In e na ional License. The
images o o he hi d pa y ma e ial in his a icle a e included in he a icle’s C ea i e
Commons license, unless indica ed o he wise in he c edi line; i he ma e ial is no included unde he
C ea i e Commons license, use s will need o ob ain pe mission om he license holde o ep oduce he
ma e ial. To iew a copy o his license, isi h p://c ea i ecommons.o g/licenses/by/4.0
Me ada a associa ed wi h his Da a Desc ip o is a ailable a h p://www.na u e.com/sda a/ and is eleased
unde he CC0 wai e o maximize euse.
© The Au ho (s) 2017
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 23
Sebas ian Beie
1,*
, Axel Himmelbach
1,*
, Ch is ian Colmsee
1
, Xiao-Qi Zhang
2
,
Robe o A. Ba e o
3
, Qisen Zhang
4
, Lin Li
5
, Micha Baye
6
, Daniel Bolse
7
, S e an Taudien
8
,
Ma co G o h
8
, Ma ius Felde
8
, Alex Has ie
9
, Hana Šimko á
10
, Helena S aňko á
10
,
Jan V ána
10
, Saki Chan
9
, Ma ía Muñoz-Ama iaín
11
, Rachid Ouni
12
, S e e Wanamake
11
,
Thomas Schmu ze
1
, Lala Aliye a-Schno
1
, S e ano G asso
13
, Jaakko Tanskanen
14
,
Dha anya Sampa h
15
, Da en Hea ens
15
, Sujie Cao
16
, B e Chapman
3
, Fei Dai
17
,
Yong Han
17
,HuaLi
16
, Xuan Li
16
, Chongyun Lin
16
, John K. McCooke
3
, Cong Tan
3
,
Songbo Wang
16
, Shuya Yin
17
, Gao eng Zhou
2
, Jesse A. Poland
18
, Ma hew I. Bellga d
3
,
And eas Houben
1
, Ja osla Doležel
10
, Sa ah Ayling
15
, S e ano Lona di
12
, Pe e Lang idge
19
,
Ga y J. Muehlbaue
5,20
, Paul Ke sey
7
, Ma hew D. Cla k
15,21
, Ma io Caccamo
15,22
, Alan
H. Schulman
14
, Ma hias Pla ze
8
, Timo hy J. Close
11
, Ma s Hansson
23
, Guoping Zhang
17
,
Ilka B aumann
24
, Chengdao Li
2,25,26
, Robbie Waugh
6,27
, Uwe Scholz
1
, Nils S ein
1,28
& Ma in Masche
1,29
1
Leibniz Ins i u e o Plan Gene ics and C op Plan Resea ch (IPK) Ga e sleben, 06466 Seeland, Ge many.
2
School o
Ve e ina y and Li e Sciences, Mu doch Uni e si y, Mu doch, Wes e n Aus alia 6150, Aus alia.
3
Cen e o Compa a i e
Genomics, Mu doch Uni e si y, Mu doch, Wes e n Aus alia 6150, Aus alia.
4
Aus alian Expo G ains Inno a ion Cen e,
Sou h Pe h, Wes e n Aus alia 6151, Aus alia.
5
Depa men o Ag onomy and Plan Gene ics, Uni e si y o Minneso a, S
Paul, Minneso a 55108, USA.
6
The James Hu on Ins i u e, Dundee DD2 5DA, UK.
7
Eu opean Molecula Biology
Labo a o y—The Eu opean Bioin o ma ics Ins i u e, Hinx on CB10 1SD, UK.
8
Leibniz Ins i u e on Aging—F i z Lipmann
Ins i u e (FLI), 07745 Jena, Ge many.
9
BioNano Genomics Inc., San Diego, Cali o nia 92121, USA.
10
Ins i u e o
Expe imen al Bo any, Cen e o he Region Haná o Bio echnological and Ag icul u al Resea ch, 78371 Olomouc, Czech
Republic.
11
Depa men o Bo any & Plan Sciences, Uni e si y o Cali o nia, Ri e side, Ri e side, Cali o nia 92521, USA.
12
Depa men o Compu e Science and Enginee ing, Uni e si y o Cali o nia, Ri e side, Ri e side, Cali o nia 92521, USA.
13
Depa men o Ag icul u al and En i onmen al Sciences, Uni e si y o Udine, 33100 Udine, I aly.
14
G een Technology,
Na u al Resou ces Ins i u e (Luke), Viikki Plan Science Cen e, and Ins i u e o Bio echnology, Uni e si y o Helsinki,
00014 Helsinki, Finland.
15
Ea lham Ins i u e, No wich NR4 7UH, UK.
16
BGI-Shenzhen, Shenzhen 518083, China.
17
College o Ag icul u e and Bio echnology, Zhejiang Uni e si y, Hangzhou 310058, China.
18
Kansas S a e Uni e si y,
Whea Gene ics Resou ce Cen e , Depa men o Plan Pa hology and Depa men o Ag onomy, Manha an, Kansas 66506,
USA.
19
School o Ag icul u e, Uni e si y o Adelaide, U b ae, Sou h Aus alia 5064, Aus alia.
20
Depa men o Plan and
Mic obial Biology, Uni e si y o Minneso a, S Paul, Minneso a 55108, USA.
21
School o En i onmen al Sciences,
Uni e si y o Eas Anglia, No wich NR4 7UH, UK.
22
Na ional Ins i u e o Ag icul u al Bo any, Camb idge CB3 0LE, UK.
23
Depa men o Biology, Lund Uni e si y, 22362 Lund, Sweden.
24
Ca lsbe g Resea ch Labo a o y, 1799 Copenhagen,
Denma k.
25
Depa men o Ag icul u e and Food, Go e nmen o Wes e n Aus alia, Sou h Pe h, Wes e n Aus alia 6150,
Aus alia.
26
Hubei Collabo a i e Inno a ion Cen e o G ain Indus y, Yang ze Uni e si y, Jingzhou, Hubei 434025, China.
27
School o Li e Sciences, Uni e si y o Dundee, Dundee DD2 5DA, UK.
28
School o Plan Biology, Uni e si y o Wes e n
Aus alia, C awley 6009, Aus alia.
29
Ge man Cen e o In eg a i e Biodi e si y Resea ch (iDi ) Halle-Jena-Leipzig, 04103
Leipzig, Ge many. *These au ho s con ibu ed equally o his wo k.
www.na u e.com/sda a/
SCIENTIFIC DATA |4:170044 |DOI: 10.1038/sda a.2017.44 24