This is a sel -a chi ed e sion o an o iginal a icle. This e sion
may di e om he o iginal in pagina ion and ypog aphic de ails.
Au ho (s):
Ti le:
Yea :
Ve sion:
Copy igh :
Righ s:
Righ s u l:
Please ci e he o iginal e sion:
CC BY-NC-ND 4.0
h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/
T ansposable elemen s in D osophila mon ana om ha sh cold en i onmen s
© The Au ho (s) 2024
Published e sion
Tahami, Mohadeseh S.; Va gas-Cha ez, Ca los; Poikela, Noo a; Co onado-
Zamo a, Ma a; González, Jose a; Kanka e, Maa ia
Tahami, M. S., Va gas-Cha ez, C., Poikela, N., Co onado-Zamo a, M., González, J., & Kanka e, M.
(2024). T ansposable elemen s in D osophila mon ana om ha sh cold en i onmen s. Mobile
DNA, 15, A icle 18. h ps://doi.o g/10.1186/s13100-024-00328-7
2024
RESEARCH Open Access
© The Au ho (s) 2024. Open Access This a icle is licensed unde a C ea i e Commons A ibu ion-NonComme cial-NoDe i a i es 4.0
In e na ional License, which pe mi s any non-comme cial use, sha ing, dis ibu ion and ep oduc ion in any medium o o ma , as long as you
gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons licence, and indica e i you modi ied he
licensed ma e ial. You do no ha e pe mission unde his licence o sha e adap ed ma e ial de i ed om his a icle o pa s o i . The images o
o he hi d pa y ma e ial in his a icle a e included in he a icle’s C ea i e Commons licence, unless indica ed o he wise in a c edi line o he
ma e ial. I ma e ial is no included in he a icle’s C ea i e Commons licence and you in ended use is no pe mi ed by s a u o y egula ion
o exceeds he pe mi ed use, you will need o ob ain pe mission di ec ly om he copy igh holde . To iew a copy o his licence, isi h p://
c ea i ecommons.o g/licenses/by-nc-nd/4.0/.
Tahami e al. Mobile DNA (2024) 15:18
h ps://doi.o g/10.1186/s13100-024-00328-7 Mobile DNA
*Co espondence:
Jose a González
[email p o ec ed]
Maa ia Kanka e
maa ia.kanka [email p o ec ed]
1Depa men o Biological and En i onmen al Science, Uni e si y o
Jy äskylä, Jy äskylä, Finland
2Ins i u e o E olu iona y Biology, CSIC, UPF, Ba celona, Spain
3Cen e o Biological Di e si y, Uni e si y o S And ews, S And ews, UK
4Ins i u Bo ànic de Ba celona (IBB), CSIC-CMCNB, Ba celona 08038,
Ca alonia, Spain
Abs ac
Backg ound Subs an ial disco e ies du ing he pas cen u y ha e e ealed ha ansposable elemen s (TEs) can
play a c ucial ole in genome e olu ion by a ec ing gene exp ession and inducing gene ic ea angemen s, among
o he molecula and s uc u al e ec s. Ye , ou knowledge on he ole o TEs in adap a ion o ex eme clima es is s ill
a i s in ancy. The a ailabili y o long- ead sequencing has opened up he possibili y o iden i y and s udy po en ial
unc ional e ec s o TEs wi h highe p ecision. In his wo k, we used D osophila mon ana as a model o cold-adap ed
o ganisms o s udy he associa ion be ween TEs and adap a ion o ha sh clima es.
Resul s Using he PacBio long- ead sequencing echnique, we de no o iden i ied and manually cu a ed TE
sequences in i e D osophila mon ana genomes om eco-geog aphically dis inc popula ions. We iden i ied 489
new TE consensus sequences which ep esen ed 92% o he o al TE consensus in D. mon ana. O e all, 11–13% o
he D. mon ana genome is occupied by TEs, which as expec ed a e non- andomly dis ibu ed ac oss he genome. We
iden i ied i e po en ially ac i e TE amilies, mos o hem om he e o ansposon class o TEs. Addi ionally, we ound
TEs p esen in he i e analyzed genomes ha we e loca ed nea by p e iously iden i ied cold ole an genes. Some
o hese TEs con ain p omo e elemen s and ansc ip ion binding si es. Finally, we de ec ed TEs nea by ixed and
polymo phic in e sion b eakpoin s.
Conclusions Ou esea ch e ealed a signi ican numbe o newly iden i ied TE consensus sequences in he
genome o D. mon ana, sugges ing ha non-model species should be s udied o ge a comp ehensi e iew o he
TE epe oi e in D osophila species and beyond. Genome anno a ions wi h he new D. mon ana lib a y allowed us o
iden i y TEs loca ed nea by cold ole an genes, and p esen a high popula ion equencies, ha con ain egula o y
egions and a e hus good candida es o play a ole in D. mon ana cold s ess esponse. Finally, ou anno a ions also
allow us o iden i y o he i s ime TEs p esen in he b eakpoin s o h ee D. mon ana in e sions.
Keywo ds D osophila mon ana, Cold adap a ion, T ansposable elemen s, Ac i e TEs, Ch omosomal in e sions
T ansposable elemen s in D osophila mon ana
om ha sh cold en i onmen s
Mohadeseh S.Tahami1, Ca losVa gas-Cha ez2, Noo aPoikela1,3, Ma aCo onado-Zamo a2,4, Jose aGonzález2,4* and
Maa iaKanka e1*
Page 2 o 21Tahami e al. Mobile DNA (2024) 15:18
Backg ound
En i onmen al s ess is one o he key ai s a ec ing
species su i al and dis ibu ion, especially in no he n
a eas whe e empe a u e luc ua ions a e less p edic -
able and ex eme wea he e en s a e becoming mo e
equen because o clima e change [1, 2]. Recen s ud-
ies ha e indica ed ha species may adap h ough a i-
ous in e plays be ween di e en gene ic elemen s in he
o ganism’s genome [3–7]. One o such elemen s d i -
ing adap a ion a e ansposable elemen s (TEs). TEs a e
epe i i e sequences wi h he abili y o eplica e hem-
sel es independen ly o he genome and change hei
posi ion he eby gene a ing di e se mu a ions. They ep-
esen a signi ican po ion o he genome in many spe-
cies, o example, o e wo hi ds o he human genome
[8] and 90% o he whea genome [9] a e TEs. In insec s,
TE con en also con ibu es signi ican ly o genome size
a ia ion [10], o example, in D osophila lies, TE con-
en can a y be ween 3% and 30% [11]. The ac i i y and
he abundance o hese highly epe i i e sequences in he
genome sugges s ha TEs could be impo an ac o s in
he e olu ion o many species.
TEs can impac species adap a ion o en i onmen al
changes. Fo example, hey a e known o be ac i a ed
unde s ess [12] and can a ec he species esponse o
en i onmen al s esso s such as hea shock [13], cold
[14], pes icides [15], insec icides [16] and oxida i e s ess
[17] hus c ea ing a ole ance agains hese s esso s.
TEs can ha e di e se e ec s on he hos ’s genome based
on he loca ion hey a e inse ed and he egula o y
sequences hey con ain, e.g. hey can c ea e inse ional
mu a ions, up- o down egula e nea by genes, o c ea e
new gene a ian s by in oducing new exons, new s op
codons, o al e na i e splice si es [18–20]. In some cases,
one o hese newly gene a ed alleles may become adap-
i e in esponse o en i onmen al s ess. This so-called
‘adap i e TE’ will be co-op ed and hus ole a ed by na u-
al selec ion. The case o he co-op ed FB i0019985 oo
solo-LTR ansposon, loca ed in he p omo e o he
Lime gene o D. melanogas e is an example in which
he adap i e TE a ec s gene exp ession associa ed wi h
cold-s ess and immune ole ance [21, 22]. In he case o
immune ole ance, he au ho s show ha FB i0019985
oo solo-LTR inse ion is adding unc ional ansc ip-
ion ac o binding si es (TFBS) o ansc ip ion ac o s
in ol ed in immune esponse [22]. I he s ess condi-
ions pe sis , co-op ed TEs migh become ixed in he
popula ion [23].
Ano he way ha TEs can pa icipa e in en i onmen al
adap a ion is by inducing ec opic ecombina ion leading
o s uc u al a ia ions such as in e sions [24, 25]. One
such case is epo ed o a Galileo ansposon ha has
gene a ed h ee polymo phic in e sions in D osophila
buzza ii h ough ec opic ecombina ion [26]. In e sions
in u n can play a ole in species adap a ion and specia-
ion in many ways [ e iewed in e.g. 27–29]. Thei main
e olu iona y signi icance is ha hey can educe ecom-
bina ion be ween a o able combina ions o alleles while
p o ec ing se s o locally adap ed genes, he eby p omo -
ing ecological di e gence, and os e ing ep oduc i e
isola ion wi hin species [30]. Fo example, in e sions a e
p e ailing d i e s o popula ion di e gence in D. i ilis
species g oup [31]. Despi e he abundance o bo h in a-
speci ic polymo phic and in e speci ic ixed in e sions
in Dip e a, s udies on he ole o TEs on ch omosomal
in e sions a e s ill sca ce.
The no he n mal ly, D osophila mon ana, is a widely
dis ibu ed insec species which belongs o he D osoph-
ila i ilis g oup [32] and is one o he mos cold- ole an
D osophila species [33, 34]. These no he n mal lies can
spend up o 6 mon hs in subze o empe a u es and se -
e al popula ions ha e adap ed o li e in la i udes abo e
he A c ic Ci cle in No he n Scandina ia. In lowe la i-
udes, lies inhabi high ele a ions (abou 3,000m) e.g.,
in he Rocky Moun ains in Colo ado (USA). They also
exis in he wa me coas al a eas in Washing on and O e-
gon (USA) and hence show a wide ange o eco-clima e
adap a ion [35]. Popula ions o D. mon ana lies a e clus-
e ed in o Eu opean, No h Ame ican and Asian popula-
ions [36] and he cu en geog aphical pa e n is likely
he esul o sp eading No h Ame ican popula ions in o
Eu asia h ough he Be ing S ai be ween 450,000 and
1,750,000 yea s ago [35] esul ing in i s unique in aspe-
ci ic di e si ica ion. In D. mon ana, adap a ion o he
ex eme clima e is modula ed by se e al genes associa ed
wi h cold ole ance [37–43]. These genes egula e a a i-
e y o unc ions om gene al cold esis ance o cu icu-
la and ol ac o y p ocesses and pho ope iodic diapause,
enhancing he species o e win e ing abili y [34, 42]. I is
also possible ha some TEs may ha e a ole in he egula-
ion o he cold- ela ed genes, leading o he species high
esis ance, e.g. o eezing empe a u es. Mo eo e , based
on ea lie poly ene ch omosome s udies, D. mon ana
ca ies se e al ixed and polymo phic in e sions [44, 45].
So a , one ixed in e sion [46] (in compa ison o D o-
sophila la omon ana), and wo polymo phic in e sions
(Poikela e al., in p ep.) has been cha ac e ized a he
genomic le el, bu he link be ween TEs and in e sions in
D. mon ana has no been s udied ye .
He e, we iden i ied TEs in D. mon ana lies ac oss
i s dis ibu ion ange o he i s ime using long- ead
genome sequence da a. To s udy he po en ial ole o TEs
in adap a ion o no he n en i onmen s, we i s p o-
duced a comp ehensi e manually cu a ed TE lib a y om
No h Ame ican, No h Eu opean and Asian popula ions
o D. mon ana. Nex , we in es iga ed he abundance,
densi y, dis ibu ion and ac i i y o TEs h oughou hese
genomes. Finally, we s udied he associa ion be ween TEs
Page 3 o 21Tahami e al. Mobile DNA (2024) 15:18
and ch omosomal in e sions, as well as he link be ween
TEs and a selec ed se o cold ole an candida e genes
ound in p e ious s udies o iden i y possible adap i e
TEs. We add essed he ollowing ques ions: (i) wha is he
genome-wide TE p o ile in D. mon ana and does i di e
om ha o o he D osophila species? (ii) Can we de ec
po en ially ac i e TEs in he D. mon ana’s genome? (iii)
Can we ind any e idence o he ole o TEs in adap a-
ion o he no he n en i onmen s? and (i ) a e he e TE
inse ions loca ed nea by in e sion b eakpoin s?
Me hods
Sample collec ions
D osophila mon ana lies we e collec ed om 13 loca-
ions in No h Ame ica (NA), No h Eu ope (NE), and
Fa Eas Asia (FE) be ween 2013 and 2021 (Fig.1 and
Supplemen a y Table S1). The collec ed emales we e
b ough o he ly labo a o y o he Uni e si y o Jy äs-
kylä, Finland, and kep in cons an ligh , 19 ˚C and ~ 60%
humidi y. The emales ha had ma ed in na u e we e
allowed o lay eggs in mal ials o se e al days. The
eme ged F1 p ogeny o each emale we e kep oge he
o p oduce he nex gene a ion and o es ablish iso emale
s ains. A e he es ablishmen o iso emale s ains, all
wild-collec ed emales and hei F1 p ogeny, excep hose
collec ed in FE, we e s o ed in 70% E OH a -20 ˚C.
Fo PacBio (Paci ic Biosciences) long- ead sequenc-
ing, D. mon ana emales we e collec ed om i e iso e-
male s ains o igina ing om i e loca ions: h ee in NA,
Sewa d (monSE13F37; Alaska, USA) and Jackson (mon-
JX13F48; Wyoming, USA) om a p e ious s udy [46]
and C es ed Bu e (mon34CC5; Colo ado, USA) om he
cu en s udy; one in NE, Oulanka (monOU13F149; Fin-
land); and one in FE, Kamcha ka (monKR1309; Russia)
(Fig.1 and Supplemen a y Tables S1 & S2). P io o col-
lec ing he lies o PacBio sequencing, iso emale s ains
we e kep in he labo a o y o ~ 50 gene a ions (Supple-
men a y Table S2).
Fo Illumina sho - ead sequencing, we used he s o ed
wild-collec ed emales o hei F1 daugh e s om 12 col-
lec ion si es in NA and NE. Because he wild-collec ed
emales o hei F1 daugh e s om FE we e no s o ed
a -20 ˚C, we pe o med Illumina sequencing on emales
collec ed om iso emale s ains kep in he labo a o y o
~ 60 gene a ions (Fig.1 and Supplemen a y Tables S1 &
S3).
Long- ead sequencing
PacBio sequencing echnique was used o he i e D.
mon ana iso emale s ains. Excep o Sewa d and Jack-
son [46], all samples we e sequenced in his s udy. Fo
all samples, DNA was ex ac ed om a pool o 60 whole
bodies o emale lies. De ails o samples, DNA ex ac-
ion, and sequencing me hods a e gi en in Supplemen-
a y Table S2.
Sho - ead sequencing
Illumina whole-genome sequencing was pe o med o
3–11 wild-collec ed o F1 emales pe popula ion om
NA and NE, and o 3 emales collec ed om FE iso e-
male s ains (Fig.1 and Supplemen a y Table S3). The
majo i y o he ly samples (94/101) we e sequenced o
his s udy, bu se en samples we e ob ained om Poikela
e al. [46] (see Supplemen a y Table S3). DNA ex ac ions
we e ca ied ou o single emales using ce yl ime h-
ylammonium b omide (CTAB) solu ion wi h RNAse
ea men , Phenol-Chlo o o m-Isoamyl alcohol (25:24:1)
and Chlo o o m-Isoamyl alcohol (24:1) washing s eps
and e hanol p ecipi a ion a he Uni e si y o Jy äskylä,
Finland. De ails o lib a y p epa a ion, sequencing ech-
nologies, loca ions, and yea s a e gi en in Addi ional ile
1, Supplemen a y Table S3. We gene a ed on a e age
19-97X co e age pe sample (Supplemen a y Table S3).
De no o genome assemblies and sca olding
We ob ained de no o genome assemblies o D. mon-
ana Sewa d and Jackson om Poikela e al. [46] and
cons uc ed new de no o genome assemblies o D. mon-
ana C es ed Bu e, Kamcha ka and Oulanka ollowing
he same p o ocol. In b ie , we assembled he genomes
using PacBio and he espec i e Illumina eads wi h
w dbg2 pipeline 2.5 [47] and MaSuRCA hyb id assem-
ble 3.3.9 [48], a e which he assembly con igui y was
imp o ed using quickme ge 0.3 [49] and he assemblies
Fig. 1 Sample loca ions o he D. mon ana popula ions analyzed in his s udy. The ed ci cles indica e he loca ion o he i e genomes ha we e used
o gene a e he o iginal TE lib a y (Supplemen a y Table S2). The ed and blue ci cles show he loca ion o popula ions used in he Illumina sequencing
(Supplemen a y Table S3)
Page 4 o 21Tahami e al. Mobile DNA (2024) 15:18
we e polished wi h he same Illumina eads using Pilon
1.23 [50]. Finally, uncollapsed he e ozygous egions
we e emo ed using pu ge_dups 1.0.1 [51] and genomic
con aminan s we e emo ed using BlobTools 1.1 [52].
We es ima ed he comple eness o he assemblies using
he BUSCO pipeline 5.1.2 based on he Dip e a da abase
“dip e a_odb10” [53], which sea ches o he p esence o
3,285 conse ed single copy Dip e a o hologues.
Ch omosome-le el genome assembly o D. mon ana
Sewa d was also ob ained om Poikela e al. [46]. This
assembly was used o sca old he o he genomes (Jack-
son, C es ed Bu e, Kamcha ka and Oulanka) using a
e e ence-genome-guided sca olding ool RagTag 2.1.0
[54] which o ien s and o de s he inpu con igs based on
a e e ence using minimap2 [55].
Building a de no o TE lib a y
The i e de no o assembled genomes we e used o de
no o iden i ica ion o TEs using he TEdeno o pipeline in
he REPET package 3.0 [56]. In sho , epea ed egions
we e de ec ed by sel -alignmen o he genomic chunks,
he TE candida es we e clus e ed using h ee me hods:
Recon 1.08 [57], G oupe 2.27 [58] and Pile 1.0 [59].
Consensus sequences we e de i ed om mul iple align-
men s in each clus e . The consensuses we e classi ied
based on s uc u al ea u es o homology ma ch wi h
he Repbase e e ence lib a y 20.05 [60] and hidden
Ma ko model (HMM) p o iles. To anno a e consensus
TEs, each TE lib a y was blas ed agains i s hos genome
using h ee di e en ools Blas e 2.25, Repea Maske
4.0.6, and CENSOR 4.1, in eg a ed in o TEanno pipe-
line [56]. To apply a s a is ical il e ing, a andomized
genomic chunk alignmen s was also applied in o de o
calcula e he high-sco ing segmen pai s (HSP). The
nex s eps il e ed and combined HSPs by keeping only
he ones ha ing a sco e highe han he HSP h eshold,
emo ed spu ious HSPs, and calcula ed iden i y pe cen -
ages. We also kep and me ged he sho simple epea
(SSR) da a by applying s ep 4 and 5 o TEanno . Then,
consensus sequences wi h less han one ull leng h copy,
iden i ied by TEanno as sequences wi h 95% o simila i y
ma ch o e he ull leng h o TE consensus, h oughou
he genome we e il e ed ou and he TEanno p ocess
was epea ed on he upda ed lib a y. Fo manual cu a-
ion, he consensus sequences o each indi idual lib a y
we e isualized in he In eg a i e Genomics Viewe (IGV)
2.8.0 [61] along wi h hei s uc u al ea u es, conse ed
domains, nucleo ide, and amino acid homology ma ches.
We cu a ed iden i ied TEs based on Wicke ’s ea u es
[62] and we ollowed se e al s eps o educe he alse
posi i e e o in ou TE lib a y and o emo e he edun-
dancy. Fo mo e de ails on he manual cu a ion p ocess
please e e o he Addi ional ile 1, Supplemen a y Table
S4 and Addi ional ile 2. To alida e ou manual cu a ion
p ocess, we un MCHelpe on he inal lib a y [63].
He e och oma in iden i ica ion
Because we we e in e es ed in TE inse ions wi h po en-
ial unc ional impac , we iden i ied he e och oma in
egions o be emo ed om u he TE analyses. A e
he manual TE cu a ion, he inal lib a y was used o
anno a e TEs ac oss each ch omosome-le el assembly
using Repea Maske 4.1.0 [64] wi h ollowing pa ame-
e s: -e ncbi -nolow -xsmall -gccalc. Since no p io in o -
ma ion on he he e och oma in egions o D. mon ana
genome was a ailable, we plo ed he o al TE densi ies
pe 5kb genomic bins. We conside ed he egions wi h
he highes TE densi y o be he e och oma ic egions.
To iew he TE densi y plo please e e o Addi ional
ile 3 Supplemen a y Figu e S1 and o he able he e o-
s. euch oma in coo dina es, e e o Addi ional ile 1,
Supplemen a y Table S5. We conside ed a egion he e o-
ch oma in i i was spanned by a leas 4 bins wi h o e
80% TE densi y. As expec ed, we ound a nega i e co -
ela ion be ween TE con en and gene densi y in he e o-
ch oma ic egions (Supplemen a y Table S6A).
Gene anno a ion
We used he Sewa d ch omosome-le el assembly as
he e e ence o genome anno a ion. The cu a ed TE
lib a y o D. mon ana was used o mask genomes using
Repea Maske 4.1.0 wi h he ollowing pa ame e s: -e
ncbi -nolow -xsmall -gccalc. The so -masked genome
was anno a ed using D. mon ana RNAseq da a (Illu-
mina T useq 150bp PE) [65]. RNA-seq eads we e i s
immed using as p 0.20.0 [66] and mapped agains
he so -masked genome using STAR 2.7.10a [67]. Gene
anno a ions we e ca ied ou wi h b ake 2.1.6 [68]
wi h de aul pa ame e s, wi h RNAseq as e idence using
Augus us 3.4.0 and GeneMa k-ET 4.48 implemen ed
pipelines [69–74]. Finally, he anno a ion was ans e ed
o he o he ch omosome-le el genomes using Li o
1.6.3 [75]. A ange o 14,504 o 14,725 genes we e ans-
e ed o each genome (Supplemen a y Table S6B). The
a e age e o a e was calcula ed based on he o al num-
be o exons ha we e co ec ly ans e ed o each gene
o be a 1.1%.
SNP calling
To examine i SNP a ia ion wi hin TE egions can e lec
he geog aphical a ia ion be ween D. mon ana genomes
ac oss i s dis ibu ion, Illumina eads we e used o ex ac
nucleo ide a ian s om he whole genome and om TE
egions o compa ison. Reads we e il e ed o adap -
e s and bases below quali y o 20 we e immed a bo h
ends using as p 0.21.0. T immed eads we e mapped o
he Sewa d ch omosome assembly (monSE13F37) using
Page 5 o 21Tahami e al. Mobile DNA (2024) 15:18
BWA mem 0.7.17 [76] wi h de aul pa ame e se ings.
PCR duplica es we e emo ed om so ed bam iles wi h
sambamba 0.7.0 [77]. Bam iles we e pa sed in o he
eebayes 1.3.6 [78] as inpu o a ian calling wi h no
popula ion p io , using wo bes SNP alleles, minimum
mapping quali y o 20, minimum co e age o 5 and he a
0.02. The VCF iles we e no malized based on he e e -
ence genome in 0.57721 [79]. Eigen ec o and eigen-
alues we e c ea ed using Plink 2.00a3 [80] a e SNPs
we e p uned o linkage disequilib ium (--indep-pai wise
50 10 0.1). We ini ially iden i ied 7,252,950 single nucleo-
ide a ian s o he whole genome and 1,773,239 single
nucleo ide a ian s wi hin he TE egions. A e he pos -
p ocessing, 6,133,447 high quali y SNPs we e eco e ed
om he whole genome, om which 6,130,880 we e
a iable wi h 6,128,313 (99.9%) biallelic and 2,567 (0.1%)
mul iallelic si es. And om he TE egion, 1,342,275 high
quali y SNPs we e eco e ed om which 1,340,930 we e
a iable wi h 1,339,585 (99.9%) biallelic and 1,345 (0.1%)
mul iallelic si es.
TE con en
The landscapes o TEs using Kimu a 2-pa ame e dis-
ance we e plo ed using pe l sc ip s in eg a ed in
Repea Maske 4.1.0 using he ch omosome assemblies.
To explo e TE abundance di e ences be ween genomes,
Chi-squa e es (χ2), using he chisq. es () unc ion in R,
was pe o med on he numbe o inse ions in each o he
genomes and he Pea son’s esiduals we e plo ed.
To al TE densi y was calcula ed be ween he e och o-
ma in and euch oma in by in e sec ing masked TEs wi h
he iden i ied eu/he e och oma in coo dina es (bed ools
.2.30.0 [81]). We me ged masked TEs ha o e lapped
in he Repea maske ’s ou pu ile i hey belonged o he
same o de , o he wise hey we e named ‘o e lap’ using a
cus om sc ip .
To calcula e TE densi y ac oss di e en genomic
egions, we used he anno a ed g ile o ex ac in ons,
exons, ups eam, downs eam and in e genic egions
only in he euch oma ic a ea. O e laps be ween egions
we e emo ed using bed ools sub ac . We calcula ed
he pe cen age o o al leng h o each o he ou TE
o de s (LTR, LINE, TIR, RC) pe genomic egion and in
euch oma in and he e och oma in, he esul s we e plo -
ed using he ggplo 2 package [82]. To es i TEs a e an-
domly dis ibu ed ega ding he loca ion o nea by genes,
we pe o med Chi-squa e es s a is ics in R using he
chisq. es () unc ion o he o al numbe o inse ions
o e he size o each compa men .
To check i he sequencing and assembly me ics
a ec ed he TE con en iden i ied in each genome, we
analyzed he co ela ion be ween hese me ics and TE
con en es ima es using a linea eg ession model.
Iden i ica ion o po en ially ac i e TEs
To ind TEs ha a e po en ially ac i e we il e ed he
anno a ion ile o TEs ha ha e a leas wo ull leng h
copies in he genome. We kep hose TEs ha had equal
o o abo e 99% o simila i y ma ch spanning a leas
50% o he o al TE consensus leng h using a cus om
bash sc ip . The inal candida es we e manually checked
o p o ein coding domains, e minal epea s, and a -
ge si e duplica ion (TSD) when applicable, ollowing
he me hod desc ibed in Va gas-Cha ez e al. [83]. To
calcula e he popula ion equency o hose ac i e TE
inse ions, we used PoPoola ionTE2 1.10.03 [84] and
TEMP2 0.1.4 pipelines [85] since di e en pipelines
migh iden i y di e en TE inse ions. The TE inse ions
ha we e de ec ed by bo h PoPoola ionTE2 and TEMP2
pipelines we e me ged using bed ools me ge (bed ools
.2.30.0) allowing 25bp dis ance (-d 25) o each inse -
ion posi ion.
Cold ole an genes
We selec ed a o al o 26 cold ole an associa ed genes
using p e ious s udies ela ed o D. mon ana cold ole -
ance and he published D. mon ana genome assembly
based on sho eads [37–43]. We included he genes in o
ou cold ole an candida e gene lis i hey we e disco -
e ed in a leas h ee di e en s udies (Supplemen a y
Table S7). To ans e he coo dina es o ou genome,
we pe o med a blas sea ch o he candida e genes in
he Sewa d ch omosome-le el assembly. Finally, he
new coo dina es o he genes we e checked o homolo-
gies in D. melanogas e genome using Flybase e sion
FB2023_01 [86].
Each gene was hen scanned up o 1.5kb up and down-
s eam o ind TEs. We ocused on TEs ha we e ound
o be p esen in all i e genomes analyzed. To check
whe he he selec ed TEs could be adding egula o y
sequences, we sea ched o ansc ip ion ac o binding
si es (TFBS) and p omo e mo i s in TEs up- o down-
s eam o he cold ole an genes. We downloaded TFBS
mo i s ela ed o s ess esponse in D osophila including
immune esponse, hea shock [87, 88], diapause [89–91]
and cold shock domain ac o s [92], om JASPAR da a-
base [93] (Supplemen a y Table S8). We used da a om
D. melanogas e when a ailable and o he wise om
Homo sapiens. TFBS we e iden i ied using he web e -
sion o FIMO [94] om he MEME SUITE [95, 96] wi h
de aul pa ame e s. The ElemeNT online ool was used
o iden i y p omo e mo i s [97, 98]. We used R package
gggenomes o isual ep esen a ion o ma ched TFBS
and p omo e s on he a o emen ioned TEs [99].
Popula ion analysis
A hund ed and one wild-caugh , indi idually sequenced
emales we e used o es ima e popula ion equencies o
Page 6 o 21Tahami e al. Mobile DNA (2024) 15:18
ac i e TEs and TEs loca ed nea by cold ole an genes
(Supplemen a y Table S3). A o al o 13 popula ions wi h
a minimum numbe o 3 indi iduals pe popula ion we e
analyzed using PoPoola ionTE2 1.10.03 [84] and TEMP2
0.1.4 [85]. Raw pai ed-end eads we e immed using
as p 0.21.0 wi h minimum quali y Ph ed sco e ≥ 20 (-q
20) and de aul pa ame e s.
PoPoola ionTE2
PoPoola ionTE2 allows he de ec ion o e e ence and de
no o TE inse ions in genomes. Following he manual,
we i s c ea ed a “TE-me ged- e e ence” o D. mon-
ana, which consis s o he espec i e masked e e ence
genome (monSE13F37) and he TE consensus sequences
gene a ed in his wo k. Nex , we c ea ed he “TE hie -
a chy ile” by using an ad hoc bash sc ip . T immed aw
eads o each sample we e mapped o he co espond-
ing TE-me ged- e e ence by using he local alignmen
algo i hm BWA bwasw . 0.7.18 ( 1243) [100]. Bo h end
eads we e mapped sepa a ely o he TE-me ged- e e -
ence, and he pai ed end in o ma ion was es o ed subse-
quen ly wi h module se2pe o PoPoola ionTE2. A ppileup
was gene a ed o each sample wi h he PoPoola ionTE2
ppileup unc ion (--map-qual 15). Finally, TE inse ions
we e iden i ied wi h unc ions iden i ySigna u es (--min-
coun 2), he inal se o inse ion pe sample was iden-
i ied wi h unc ion pai upSigna u es and ou pu s we e
in e sec ed wi h TE coo dina es loca ed nea by cold ol-
e an genes using bed ools in e sec .
TEMP2
To de ec de no o TE inse ions, we used he module
inse ion o TEMP2. We ook he Repea Maske anno a-
ion ha was un o PoPoola ionTE2 and ans o med
i o bed o ma using g 2bed implemen ed in BEDOPS
.2.4.39 [101]. Following he TEMP2 manual, we used
BWA mem wi h op ions -Y and -T20 o map he pai ed-
end eads o he co esponding e e ence genome. To
ge he agmen leng h o he pai ed-end eads, we used
Pica d’s Collec Inse SizeMe ics module .2.26.11 [102]
and inse ed he mean inse size o each sample. Then,
we used TEMP2 inse ion module wi h pa ame e -m 5
(pe cen age o misma ch allowed when mapping o TEs)
o de ec TE inse ions.
To de ec e e ence inse ions, we used he absence
module which anno a es he absence o e e ence TE
copies in he samples (pa ame e -x 30, he minimum
sco e di e ence be ween he bes hi and he second bes
hi o conside ing a ead as uniquely mapped). We con-
side ed inse ions ha a e in he e e ence genome and
no anno a ed by he absence module as p esen in ou
samples. We iden i ied hese inse ions using bed ools
in e sec wi h - op ion. Finally, we joined he e e -
ence inse ions wi h he ones de ec ed by he inse ion
module. Ou pu s we e in e sec ed wi h TE coo di-
na es loca ed nea by cold ole an genes using bed ools
in e sec .
Combining PoPoola ionTE2 and TEMP2 TE inse ions
To ind a eliable se o TEs, we combined he TE inse -
ions de ec ed by bo h so wa es. We used bed ools wi h
op ions me ge -i and -d 25 o collapse inse ions o e lap-
ping o allowing a maximum dis ance o 25bp be ween
he wo coo dina es in o a single call. We combined
inse ions ha belonged o he same TE amily.
Ch omosomal in e sions
B eakpoin s o one ixed and wo polymo phic D. mon-
ana in e sions we e ob ained om Poikela e al. [46] and
Poikela e al. (in p ep.) (Supplemen a y Table S9). B ie ly,
bo h long- and sho - eads we e used o iden i y he
in e sion b eakpoin s. PacBio and Illumina eads we e
mapped agains Sewa d genome assembly (monSE13F37)
using ngml 0.2.7 [103] and BWA mem, espec i ely, and
he esul ing bam iles we e pa sed in o Sni les 2.0.7
[103] and delly 1.1.6 [104] s uc u al a ian iden i i-
ca ion p og ams, espec i ely. SURVIVOR 1.0.6 [105]
was used o iden i y in e sions ha we e sha ed by bo h
s uc u al a ian iden i ica ion p og ams. Finally, he
pu a i e b eakpoin s o all in e sions we e con i med
isually wi h he IGV using bo h long- and sho - ead
da a.
To check whe he TEs we e p esen in he in e sion
b eakpoin , 50kb egions lanking each side o he b eak-
poin s we e mapped agains all i e sca old assemblies
using minimap2 o iden i y he b eakpoin coo dina es in
each assembly. To alida e he p esence/absence o each
b eakpoin , we iden i ied long eads spanning he b eak-
poin s (3kb on each side o he b eakpoin o he wo
la ge in e sions and 500bp on each side o he sho es
in e sion) using IGV. The p oximal and dis al posi ions
a e based on hei dis ance o he cen ome e [46].
Resul s
Th ee new genome assemblies om D. mon ana eco-
geog aphically dis inc popula ions
Th ee new D. mon ana e e ence genomes om h ee
clima ically and geog aphically di e gen popula ions,
C es ed Bu e in No h Ame ica (NA), Oulanka in No h-
e n Eu ope (NE) and Kamcha ka in Fa Eas Asia (FE),
we e gene a ed in his s udy (Fig.1; Table1, and Supple-
men a y Table S2). We also used he wo o he a ailable
genome assemblies, Sewa d and Jackson om No h
Ame ica D. mon ana popula ions [46]. As sequencing
echnology g ea ly impac s TE de ec ion in he genome,
we ook ad an age o long- ead sequencing echnique o
gene a e hese new assemblies and hus e ie e highe
numbe s o TEs compa ed o hose iden i ied based only
Page 7 o 21Tahami e al. Mobile DNA (2024) 15:18
on sho - eads [106]. All i e genomes we e sequenced
using bo h PacBio and Illumina echnology and assem-
bled using he same me hod. In No h Ame ica, he
popula ions o Jackson (monJX13F48) and C es ed Bu e
(mon34CC5) a e adap ed o ela i ely high la i udes and
high ele a ions (1857m and 2960m espec i ely) in he
moun ainous egion, whe eas Sewa d (monSE13F37) is a
high-la i ude coas al popula ion (35m). In No h Eu ope,
Oulanka popula ion (monOU13F149) is loca ed in no h-
e n Finland, abo e he a c ic ci cle, and is adap ed o cold
and da k win e s and sho summe s. The Asian sample
(monKR1309) is om Kamcha ka Peninsula which is a
moun ainous olcanic egion wi h long cold win e s in
Russia (Fig.1).
Upon de no o assembly, he genome sizes o D. mon-
ana anged om 175 o 184 Mb. By using a hyb id
assembly s a egy, we iden i ied mo e han 95% comple e
single copy BUSCOs, wi h N50 alues a ying be ween
0.2 and 11.0Mb (Table1). A e sca olding, he o al
genome size and BUSCO alues dec eased (90.6–93.1%,
145–151 Mb) compa ed o he non-sca olded assem-
blies. This is because we we e unable o assign all con igs
o ch omosome sca olds and he unassigned con igs
we e excluded om he inal ch omosome-le el D.
mon ana genome [46]. The de ailed in o ma ion on aw
eads, con igs and sca olds s a is ics a e gi en in Table1.
We used comple e genome assembly o TE iden i ica ion
and ch omosome le el assemblies excluding he he e o-
ch oma ic egions o he es o ou analyses in his wo k
(see Me hods).
O e 90% o he TE consensus sequences iden i ied in D.
mon ana belong o new amilies o new sub amilies
By using genomes om clima ically and geog aphically
di e ged D. mon ana popula ions, we wan ed o include
as much in aspeci ic di e si y in o ou comp ehensi e
TE lib a y as possible. We used REPET o de no o anno-
a e TEs in each o he i e genomes. On a e age, 3,221
TE consensus we e buil pe genome, and 303 consensus
pe genome we e kep a e manual cu a ion (Supple-
men a y Table S10). A e clus e ing indi idual lib a ies,
a o al o 555 non- edundan TE consensus sequences
we e eco e ed (TE lib a y is p o ided as Addi ional
ile 4). 92% o he TE consensuses a e desc ibed as new
Table 1 Raw ead and assembly s a is ics o i e D. mon ana genomes gene a ed by PacBio sequencing. N50 aw eads = he ead
leng h a which hal o he bases a e in eads longe han o equal o his alue. Raw eads me ics o Sewa d and Jackson a e aken
om Poikela e al. [46]
Popula ion C es ed Bu e Kamcha ka Oulanka Sewa d Jackson
assembly name mon34CC5 monKR1309 monOU13F149 monSE13F37 monJX13F48
Raw eads
Raw da a (Gb) 7.2 5.5 8.3 9.9 13.9
Max ead leng h (bp) 145,085 132,961 118,876 168,303 111,203
A e age ead leng h (bp) 5607.9 6268.7 5830 5905.9 6422.3
Mean co e age 39.9 30.6 46.4 54.8 77.1
N50 aw eads 10,260 9233 8836 11,366 8045
De no oassemblies
Assembly leng h (Mb) 175.0 176.7 178.9 184.3 181.0
Numbe o con igs 739 1675 1097 324 796
Longes con ig (Mb) 4.5 1.2 3.7 29.1 15.2
N50 (Mb) 0.8 0.2 0.5 11.0 1.3
N50 coun 55 253 97 5 20
GC le el (%) 0.402 0.401 0.401 0.402 0.402
To al comple e BUSCOs (%; n: 3285) 97.2 96 96.4 98.1 98.5
Single copy BUSCOs (%) 96.7 95.6 96.0 97.6 98.0
Duplica ed BUSCOs (%) 0.5 0.4 0.4 0.5 0.5
Sca old assemblies
Assembly leng h (Mb) 148.2 148.7 149.8 145.5 151.8
Numbe o sca olds 6 6 6 6 6
Sca old leng h (Mb) 34.8 34.4 33.3 32.5 33.9
N50 (Mb) 26.6 26.5 26.7 26.5 28.4
N50 coun 3 3 3 3 3
GC le el (%) 0.402 0.402 0.401 0.403 0.402
To al comple e BUSCOs (%; n: 3285) 92.4 90.6 90.6 91.9 93.1
Single copy BUSCOs (%) 92.2 90.4 90.4 91.7 92.8
Duplica ed BUSCOs (%) 0.2 0.2 0.2 0.20 0.3
Accession numbe SAMN27782986 SAMN27782984 SAMN27782985 SAMN27782981 SAMN27782982
Page 8 o 21Tahami e al. Mobile DNA (2024) 15:18
amilies (445) and new sub amilies (66) and only he
emaining 8% co esponds o p e iously known ami-
lies (44) (Fig. 2 and Supplemen a y Table S11). Long
e minal epea s (LTR) a e he mos abundan TE o de
in he D. mon ana lib a y o which, Gypsy supe amily
is he mos abundan (80%) ollowed by Bel-Pao (14%).
LINE and DNA ansposons (TIR, e minal in e ed
epea , and RC, olling ci cle elemen s) con ibu e almos
equally o he TE lib a y (12%) wi h he highes abun-
dance o Jockey (68%) and Tc1-Ma ine (40%) in each
o de , espec i ely (Fig.2). We used he manual cu a ed
lib a y o anno a e each o he i e e e ence genomes
and analyzed he abundance, dis ibu ion, and ac i i y
o TE sequences. We i s analyzed whe he di e ences
in sequencing o assembly me ics explained a ia ion
in he TE me ics es ima ed (Supplemen a y Table S12).
The N50 o con igs had a signi ican e ec on he o al
numbe o aw TE sequences disco e ed in he he e o-
ch oma in, while sca old s a is ics did no impac any
o he TE me ics analyzed (Supplemen a y Table S12).
Because all o he ollowing analyses ha e been made
on he euch oma in, ou esul s a e hus no a ec ed by
sequencing o assembly di e ences ac oss genomes.
We also explo ed he gene ic a ia ion ac oss he whole
genome and wi hin TE sequences in he i e D. mon ana
samples by pe o ming a p incipal componen analysis
(PCA) wi h biallelic SNPs. Resul s om he wo PCA
analyses a e highly conco dan , wi h he i s PC sepa-
a ing he h ee NA popula ions om he Oulanka (NE)
and Kamcha ka (FE) popula ions and explaining simila
amoun s o a ia ion: 33 and 32% o he whole genome
and TE egion SNPs, espec i ely (Fig.3A and B and Sup-
plemen a y Table S13). The second PC, which explains
28 and 29% o he a ia ion, sepa a es he C es ed Bu e
popula ion (NA) om he o he wo NA popula ions:
Jackson and Sewa d. This pa e n likely e lec s he
demog aphy o hese popula ions (Poikela e al., in p ep)
and shows he high le el o in aspeci ic a ia ion in he
D. mon ana genomes analyzed.
T ansposable elemen s con ibu e 11–13% o he D.
mon ana genomes
Whole genome TE con en ac oss genomes a ied
be ween 10.7% and 13% (Table2). In all i e genomes,
LTRs we e he mos abundan o de , ollowed by olling
ci cles (RC), LINEs, and e minal in e ed epea s (TIRs)
(Fig. 4A). TE abundance a he o de and supe amily
le el a ied ac oss genomes (Fig.4B and C). In Sewa d,
LTR (χ2 esidue − 8.2, p- alue < 0.05) is he leas abundan
and RC (χ2 esidue 5.7, p- alue < 0.05) is he mos abun-
dan o de compa ed wi h he o he genomes analyzed.
A he supe amily le el, Sewa d shows he highes a ia-
ion in TE con en ; Gypsy and Bel-Pao ha e less inse -
ions compa ed wi h o he genomes (χ2 esidue − 11.1
Fig. 2 Consensus sequence classi ica ions o he cu a ed TE lib a y buil om D. mon ana genomes. (A) Pe cen age o new o known TE consensus (cons)
sequences. (B) Numbe s o consensus elemen s classi ied as LTR, DNA (RC and TIR) and LINEs
Page 15 o 21Tahami e al. Mobile DNA (2024) 15:18
Table 3 T ansposable elemen s ha we e ound in he b eakpoin egions (3kb/500bp) o he h ee in e sions analyzed in his wo k.
TEs ha a e p esen only in he genomes wi h in e sion b eakpoin s (and no in o he genomes) a e in bold
Ch In e sion coo dina es Genomes TE con en
P oximal Dis al P oximal b eakpoin Dis al b eakpoin
X 11,174,369–11,174,380 3,994,112–3,994,127 Sewa d R1-2_Dmon
Ma e ick-1_Dmon
TART_DV-Dmon-B
Heli on-5_Dmon
Gypsy-102_Dmon
Heli on-10_Dmon
Heli on-2N1_D i -Dmon-B
Heli on-2N1_D i -Dmon-A
Heli on-1N1_D i
11,136,142–11,136,155 3,936,207–3,936,246 Jackson R1-2_Dmon
Heli on-5_Dmon
Gypsy-102_Dmon
Heli on-10_Dmon
Heli on-5_Dmon
Heli on-2N1_D i -Dmon-B
Heli on-2N1_D i -Dmon-A
Heli on-1N1_D i
11,110,783–11,110,796 3,984,908–3,984,929 Kamcha ka R1-2_Dmon
Ma e ick-1_Dmon
Heli on-5_Dmon
Gypsy-102_Dmon
Heli on-10_Dmon
Heli on-5_Dmon
Heli on-2N1_D i -Dmon-B
Heli on-2N1_D i -Dmon-A
Heli on-1N1_D i
11,304,271–11,304,284 4,034,972–4,035,011 Oulanka Ma e ick-1_Dmon
Heli on-5_Dmon
Gypsy-102_Dmon
Heli on-10_Dmon
Heli on-8_Dmon
Heli on-2N1_D i -Dmon-B
Heli on-2N1_D i -Dmon-A
Heli on-1N1_D i
Ma e ick-1_Dmon
Ma e ick-1_Dmon
10,289,870–10,289,883 3,549,579–3,549,600 C es ed Bu e R1-2_Dmon
Heli on-5_Dmon
Gypsy-102_Dmon
Heli on-10_Dmon
Heli on-5_Dmon
Heli on-1N1_D i
Heli on-2N1_D i -Dmon-B
Heli on-2N1_D i -Dmon-A
Heli on-1N1_D i
Ma e ick-1_Dmon
4 6,596,948–6,599,447 7,298,775 − 7,298,610 C es ed Bu e Gypsy-5_DVi -I-Dmon-F
Gypsy-40_Dmon
Gypsy-5_DVi -I-Dmon-A
MITE-1_Dmon
MITE-1_Dmon
Heli on-1_D i -Dmon-C
Gypsy-215_Dmon
Gypsy-181_Dmon
4 2,906,226 2,907,786–2,907,896 Oulanka Heli on-4_Dmon
Gypsy-83_Dmon
Gypsy-70_Dmon
Heli on-5_Dmon
Gypsy-8_DVi -I-Dmon-B
ULYSSES_I-Dmon-H
Heli on-4_Dmon
Heli on-1_D i -Dmon-C
3,165,124–3,165,142 3,166,689 Kamcha ka Heli on-4_Dmon
Gypsy-83_Dmon
Gypsy-70_Dmon
Heli on-5_Dmon
Gypsy-8_DVi -I-Dmon-B
Heli on-4_Dmon
Page 16 o 21Tahami e al. Mobile DNA (2024) 15:18
o de and Gypsy supe amily (Fig. 2). Only 8% o he
TE consensus sequences co esponded o known TE
amilies, mos ly o D. i ilis anno a ions, which is mo e
closely ela ed o he s udied species han D. melanogas-
e (Supplemen a y Table S11). The e a e se e al po en-
ial easons o he high numbe o new TE amilies.
P obably he mos impo an eason is ha D. mon ana
is highly di e ged om he o he model and non-model
D osophila species o which TE lib a ies exis in da a-
bases (be ween ~ 40 MYA om D. melanogas e and 9
MYA om D. i ilis) [112, 113]. Ano he plausible eason
comes om he species’ unique dis ibu ion in he Hol-
a c ic egion. Upon coloniza ion in o new habi a s, spe-
cies can be exposed o he in asion o new TEs h ough
ho izon al ans e om new pa asi es and om o he
conspecies [114]. Simila e en s could ha e happened in
D. mon ana while occupying new cold habi a s.
Ou p incipal componen analysis (PCA) on genome-
wide dis ibu ion o SNPs and wi hin TE egions shows
he high le el o in aspeci ic a ia ion ha we aimed
o include in o ou TE lib a y (Fig.3). The PCA depic s
he high di e gence no only be ween D. mon ana
popula ions om No h Ame ica and No he n Eu ope,
bu also among No h Ame ican popula ions as C es ed
Bu e (Colo ado) is highly di e ged om Jackson (Wyo-
ming) and Sewa d (Alaska). High di e gence o C es ed
Bu e and o he No h Ame ican popula ions is likely a
consequence o pas ounde e ec s po en ially associ-
a ed wi h gene ic d i , a iable selec ion p essu es, and
local adap a ion [43]. Mo eo e , he la ge in e sion on
ch omosome 4 is uniquely ixed in C es ed Bu e (Poikela
e al., in p ep.), which u he educes gene ic exchange
be ween popula ions and esul s in ele a ed gene ic
di e gence (Poikela e al., in p ep. and [27]). O e all,
en i onmen al he e ogenei y h ough ime and space can
ac as a selec i e o ce d i ing adap i e di e en ia ion
be ween popula ions. In e es ingly, Kamcha ka shows
high simila i y o he Finnish popula ion om Oulanka
ega dless o hei geog aphic dis ance. This a ini y could
be explained by he ances al ou e o expansion om
NA popula ions owa ds Eu asia h ough he Be ing
S ai [35], he e o e, he gene ic di e gence o Eu opean
and Asian popula ions would be low.
Fig. 8 Polymo phic and ixed in e sion b eakpoin s in D. mon ana. Fo each in e sion bo h b eakpoin s, p oximal (close o he cen ome e) and dis al
( a he om he cen ome e), plus ~ 3kb o 500bp on each side a e shown. When he posi ion o a b eakpoin was no iden i ied a he single base pai
le el, he in e al whe e he b eakpoin is p edic ed o be is shown in a g ey box. Genes a e shown as da k blue boxes, and TEs a e colo ed based on hei
supe amily le el. B eakpoin posi ions a e d awn based on hei same posi ion on he aligned ch omosomes in he sense di ec ion
Page 17 o 21Tahami e al. Mobile DNA (2024) 15:18
TE con en in D. mon ana
Ou de no o TE anno a ions con i med ha he amoun
o TE con en signi ican ly con ibu es o he D. mon ana
genome size as epo ed o o he species [10, 115]. The
o e all a e age TE con en in D. mon ana, 11–13%, is
simila o ha o i s close ela i es, D. i ilis (15%) [11]
and D. incomp a (13–14%) [107], and smalle han hose
o D. melanogas e (20%) [115] and D. suzukii (30%) [11].
LTRs ep esen he highes genomic ac ion o ~ 5% ol-
lowed by RC elemen s (Heli on and Ma e ick), ~ 4%,
LINE, ~ 2%, and inally TIR wi h < 1% (Fig.4A). Acco d-
ing o his obse ed pa e n, LTRs a e he mos abundan
elemen s which is ypical o many D osophila lies, how-
e e , he abundance does no ollow he assumed global
pa e n: LTRs > LINEs > TIRs > OTHERs [10, 11]. The
highe Heli on ac ion in D. mon ana genome (com-
pa ed o he global pa e n) combined wi h i s a he
pe sis en ac i i y spanned o e an ex ended cou se o
ime (Fig.6A) sugges s ha Heli ons ha e been ole -
a ed mo e han o he elemen s p obably due o hei
neu al o adap i e e ec s. The elaxed selec i e p essu e
on Heli on ansposi ion could come om he elemen ’s
own na u e and i s endency o inse nea less c i ical
genes. Howe e , esul s om phylogene ically ela i ely
close ela i es o D. mon ana such as D. buzza ii and D.
moja ensis, ha e indica ed he abundance o Heli on in
hei genomes, con ibu ing o he hypo hesis ha Heli-
ons can ha e a wide dis ibu ion in he subgenus D o-
sophila [116].
In D. mon ana a signi ican ly highe numbe o TEs
we e accumula ed in he he e och oma in egion. The
he e och oma in’s esilience o TE accumula ion is due
o he sca ci y o ac i e genes, and as expec ed, TEs we e
less ole a ed in he euch oma in (Fig.5A). Acco ding o
he D osophila 12 Genomes Conso ium (2007), 1–9% o
he euch oma in is occupied by TEs ac oss D osophila
species, and hence, he calcula ed 7% o D. mon ana is
wi hin he epo ed ange. Analysis o he TE con en
dis ibu ion be ween di e en genomic egions e eals
highe accumula ion o TEs in he in e genic a ea ol-
lowed by in ons, wi h he leas occupancy in exons
(Fig. 5D). O e all, TE dis ibu ion was no andom
h oughou he genome o be ween ch omosomes sug-
ges ing ha pu i ying selec ion ac s agains TE inse ions
due o hei mul iple dele e ious e ec s [117]. TE den-
si y was signi ican ly highe in ch omosome 4 and X and
he lowes in he igh a m o ch omosome 2 (2R). Apa
om he ch omosome sizes, he highe acquisi ion o
TEs on sex ch omosome could be associa ed wi h dosage
compensa ion [118]. Mo eo e , Wibe g e al. [43] ound
ha he op candida e SNPs o cold adap a ion in D.
mon ana a e en iched on ch omosomes 4 and X which
also include se e al in e sions (see below).
Acco ding o he landscape o TEs, all i e genomes
show simila dynamics o he TE con en . We ound ha
he D. mon ana TE landscape is compa able o ha o D.
melanogas e which also has a la ge ac ion o young and
ac i e LTR elemen s, while i di e s om he D. simulans
landscape which ins ead has a la ge ac ion o old and
deg aded elemen s [11]. The pa e n is bimodal wi h a
gene al end owa ds a small numbe o highly di e ged/
old elemen s (Fig.6A). The dynamic pa e n o TE land-
scapes is o en a e lec ion o he species demog aphic
his o y [119, 120]. Fo example, occupying new en i on-
men s will pose a ious ypes o s ess o he species o
which, cold is one s ess ac o in he no he n egions,
which can po en ially e-ac i a e mobile elemen s [12,
121]. The bimodal pa e n indica es wo bu s s o ans-
posi ion e en s ha ha e in e up ed he ansposi ion/
excision equilib ium (s anda d bell shape pa e n). This
pa e n is gene ally seen in specialis species and spe-
cies wi h small e ec i e popula ion size which happens
when he dis ibu ion is pa chy and limi ed by en i on-
men al esou ces [107]. The i e genomes also ha e simi-
la numbe s o po en ially ac i e TEs (Fig.6B), which a e
mos ly p esen a low popula ions equencies (Fig.6C)
as expec ed i hese a e young inse ions.
TEs associa ed wi h cold ole an genes
Ou analysis o he TEs loca ed nea by cold- ole an
genes in he i e genomes sugges ha some o hem
migh be playing a ole in he egula ion o hese genes
as hey a e p esen a high popula ion equencies and
con ain egula o y sequences (Fig. 7 and Supplemen-
a y Table S21). Th ee ypes o FOXO1 and one BEAF-
32 TFBS, among o he TFBS, we e ound inside some o
hese TEs (Fig.7). Heli on-2N1_D i -Dmon-B loca ed
downs eam o Yp2 and ups eam o CG12057 also has
p omo e mo i s and FOXO ansc ip ion ac o s. FOXO
ansc ip ion ac o s egula e cellula homeos asis, lon-
ge i y and esponse o s ess [122]. In D. mon ana, ac i-
a ion o FOXO has been sugges ed o be connec ed
wi h he lies’ o e win e ing abili y [89]. Mo eo e , Heli-
on-1_D i -Dmon-C ups eam o CG17571 has a ma ch
o p omo e elemen s only in he Kamcha ka genome.
This TE ma ches BEAF-32 TFBS which is a ch oma in
insula o and a ec s he egula ion o gene ansc ip-
ion [88]. BEAF-32 is also connec ed o he de elopmen
o pho o ecep o di e en ia ion du ing emb yogenesis
[91]. The e o e, hese candida e TEs ha e he po en ial o
a ec he egula ion o he lanking cold ole an genes,
as has been shown be o e o D. melanogas e TE inse -
ions [87].
TEs associa ed wi h in e sions
TEs and o he epe i i e egions ha e been sugges ed
o be in ol ed in he o igin o ch omosomal in e sions
Page 18 o 21Tahami e al. Mobile DNA (2024) 15:18
[123]. By s udying h ee in e sions ecen ly cha ac e ized
in D. mon ana ac oss i s dis ibu ion ( [46] and Poikela
e al., in p ep), we we e able o sea ch o he p esence
o TEs nea in e sion b eakpoin s. These in e sions we e
ound on ch omosomes X and 4, which also ha e he
highes TE densi y. Pe haps no su p isingly, we ound
TEs nea by all he b eakpoin s. The b eakpoin s o he
la ge ch omosome 4 in e sion ha is ixed in C es ed
Bu e ha bo s i e TE inse ions ha a e exclusi e o
C es ed Bu e and no ound in he o he popula ions.
Simila ly, he b eakpoin s o he small ch omosome 4
in e sion ound in Oulanka and Kamcha ka show TE
inse ions ha a e speci ic o hese popula ions bu a e
no ound in he o he s. These indings could sugges
ha some o hese TE inse ions ound a ound he in e -
sion b eakpoin s may ha e ac ed as subs a es o ec opic
ecombina ion and acili a ed he o igin o he in e sions,
al hough u he analysis would be needed o con i m
i [123]. Finally, he b eakpoin s o he X in e sion ha e
se e al TE inse ions ha a e p esen in all D. mon ana
popula ions, in acco dance wi h he in e sion being
ixed in D. mon ana. Some o hese inse ions could ha e
po en ially played a ole in he es ablishmen o ha X
in e sion ~ 3 MY [46], bu would equi e an in-dep h
in es iga ion in o he TEs be ween D. mon ana and
o he species o he mon ana phylad (D. la omon ana,
D. bo ealis, D. lacicola). O e all, hese TE inse ions can
g ea ly impac he o ma ion o la ge ea angemen s,
po en ially leading o signi ican e olu iona y ou comes.
Fo example, he la ge ch omosome 4 in e sion is asso-
cia ed wi h SNPs linked o cold/clima e adap a ion in D.
mon ana and wi h a educ ion in gene exchange be ween
D. mon ana popula ions (Poikela e al., in p ep.). Addi-
ionally, he X in e sion is linked o he specia ion p o-
cess be ween D. mon ana and D. la omon ana [46, 124].
Conclusions
S udying TEs in non-model species expands ou unde -
s anding o TE dynamics h oughou zoo axa, and hei
a ious impac s on he genome e olu ion. Adding ca e-
ully cu a ed and a non- edundan TE lib a y om non-
model species o he public eposi o ies signi ican ly
con ibu es o his mission. We eco e ed 555 cu a ed
non- edundan TE consensuses wi h ema kable num-
be s o new amilies and sub amilies comp ising o e 90%
o he TE lib a y in D. mon ana, using gene ically dis-
an in aspecies samples. TE con ibu ion o he whole
genome con en is gene ally smalle in D. mon ana han
i s ela i e D osophila species assuming ha di e ences
in he quali y o he genome assemblies be ween di e -
en species a e no he case. As expec ed TEs we e non-
andomly dis ibu ed, we also ound a highe densi y in
he e och oma ic egions, and he X ch omosome com-
pa ed o au osomes. Howe e , he TE abundance a he
o de le el seems o a y ac oss di e en species which
equi es mo e in es iga ion om he e olu iona y s and-
poin . Mo eo e , we ound TEs nea by o in he i s
in on o a selec ed se o p e iously iden i ied cold ol-
e an genes which ha e ansc ip ion ac o binding si es
and/o p omo e elemen s, indica ing hei po en ial
o in luencing gene’s egula ion. Howe e , ou cu en
a emp was only di ec ed o a se o genes iden i ied in
ou p e ious s udies. Undoub edly, u u e esea ch will
disco e mo e cold ela ed genes and TEs connec ed wi h
hem. Finally, o he i s ime, we ound se e al TEs
a ound he h ee in e sion b eakpoin s in D. mon ana,
which a e wo h u he s udies on hei po en ial impac
on in e sions.
Abb e ia ions
TE T ansposable elemen
NA No h Ame ica
NE No h Eu ope
FE Fa Eas
TFBS T ansc ip ion ac o binding si es
LTR Long e minal epea s
LINE Long in e spe sed nuclea elemen s
TIR Te minal in e ed epea s
RC Rolling ci cles
Supplemen a y In o ma ion
The online e sion con ains supplemen a y ma e ial a ailable a h ps://doi.
o g/10.1186/s13100-024-00328-7.
Supplemen a y Ma e ial 1
Supplemen a y Ma e ial 2
Supplemen a y Ma e ial 3
Supplemen a y Ma e ial 4
Acknowledgemen s
The au ho s would like o hank MSc Sa a Lommi o he wo k in he
labo a o y and in he ly s ain main enance. Analyses we e ca ied ou using
CSC (Finnish IT Cen e o Science) se ices.
Au ho con ibu ions
MST: W o e he manusc ip , p epa ed he ables and igu es, de eloped code
and ca ied ou he analyses, CVC: Con ibu ed o analyses and p oo eading.
NP: Con ibu ed o analyses, w i ing, and p oo eading. MCZ: Con ibu ed o
analyses, and igu es, and de eloped code, p oo eading. JG: Con ibu ed o
concep ualiza ion, w i ing, and p oo eading. MK: concei ed and designed he
p ojec , con ibu ed o w i ing, p oo eading, and p ojec managemen . All
he au ho s ead and app o ed he inal manusc ip .
Funding
JG is unded by g an PID2020-115874GB-I00 awa ded by MICIU/AEI/h ps://
doi.o g/10.13039/501100011033/ and om g an 2021 SGR 00417 awa ded
by Depa amen de Rece ca i Uni e si a s, Gene ali a de Ca alunya awa ded
o J.G. MK was unded by he Academy o Finland p ojec 322980. The unde s
had no ole in he p epa a ion and publica ion o his manusc ip . The es o
he au ho s ha e decla ed no unding associa ed wi h his wo k.
Da a a ailabili y
The da ase s suppo ing he conclusions o his a icle a e included wi hin
he a icle and i s addi ional iles. The genome assemblies a e e ie able wi h
p ojec name PRJNA828433 and p e iously published assemblies and PacBio
and Illumina aw eads a e a ailable unde Biop ojec PRJNA939085. The TE
Page 19 o 21Tahami e al. Mobile DNA (2024) 15:18
lib a y is p o ided as Addi ional ile 4. The sc ip s de eloped o he pu pose
o his s udy a e a ailable in he Gi hub eposi o y, h ps://gi hub.com/
ahami-ms/TE-p ojec .
Decla a ions
E hics app o al and consen o pa icipa e
No applicable.
Consen o publica ion
No applicable.
Compe ing in e es s
The au ho s decla e no compe ing in e es s.
Recei ed: 9 Ap il 2024 / Accep ed: 17 Sep embe 2024
Re e ences
1. Cohen J, Agel L, Ba low M, Ga inkel CI, Whi e I. Linking A c ic a iabili y
and change wi h ex eme win e wea he in he Uni ed S a es. Sci (1979).
2021;373:1116–21.
2. Coumou D, Rahms o S. A decade o wea he ex emes. Na Clim Chang.
2012;2:491–6.
3. Sp ingael D, Top EM. Ho izon al gene ans e and mic obial adap a ion o
xenobio ics: new ypes o mobile gene ic elemen s and lessons om eco-
logical s udies. T ends Mic obiol. 2004;12:53–8.
4. Ki kpa ick M, Ba on N. Ch omosome in e sions, local adap a ion and specia-
ion. Gene ics. 2006;173:419–34.
5. Sch ade L, Kim JW, Ence D, Zimin A, Klein A, Wysche zki K, e al. T ansposable
elemen islands acili a e adap a ion o no el en i onmen s in an in asi e
species. Na Commun. 2014;5:1–10.
6. Dudaniec RY, Yong CJ, Lancas e LT, S ensson EI, Hansson B. Signa u es o
local adap a ion along en i onmen al g adien s in a ange-expanding dam-
sel ly (Ischnu a Elegans). Mol Ecol. 2018.
7. Cayuela H, Do an Y, Mé o C, Lapo e M, No mandeau E, Gagnon-Ha ey
S, e al. The mal adap a ion a he han demog aphic his o y d i es gene ic
s uc u e in e ed by copy numbe a ian s in a ma ine ish. Mol Ecol.
2021;30:1624–41.
8. de Koning APJ, Gu W, Cas oe TA, Ba ze MA, Pollock DD. Repe i i e ele-
men s may comp ise o e wo- hi ds o he human genome. PLoS Gene .
2011;7:e1002384.
9. Vicien CM. T ansc ip ional ac i i y o ansposable elemen s in maize. BMC
Genomics. 2010;11:1–10.
10. Pe e sen M, A misén D, Gibbs RA, He ing L, Khila A, Maye G, e al. Di e si y
and e olu ion o he ansposable elemen epe oi e in a h opods wi h
pa icula e e ence o insec s. BMC Ecol E ol. 2019;19:1–15.
11. Mé el V, Boules eix M, Fable M, Viei a C. T ansposable elemen s in D osophila.
Mob DNA. 2020;11:1–20.
12. Ho á h V, Me enciano M, González J. Re isi ing he ela ionship be ween
ansposable elemen s and he euka yo ic s ess esponse. T ends Gene .
2017;33:832–41.
13. Maside X, Ba olomé C, Cha leswo h B. S-elemen inse ions a e associa ed
wi h he e olu ion o he Hsp70 genes in D osophila melanogas e . Cu Biol.
2002;12:1686–91.
14. Nai o K, Zhang F, Tsukiyama T, Sai o H, Hancock CN, Richa dson AO, e al.
Unexpec ed consequences o a sudden and massi e ansposon ampli ica-
ion on ice gene exp ession. Na u e. 2009;461:1130–4.
15. Amine zach YT, Macphe son JM, Pe o DA. Pes icide esis ance ia ans-
posi ion-media ed adap i e gene unca ion in D osophila. Science (1979).
2005;309:764–7.
16. Salces-O iz J, Va gas-Cha ez C, Guio L, Rech GE, González J. T ansposable
elemen s con ibu e o he genomic esponse o insec icides in D osophila
melanogas e . Philosophical T ans Royal Soc B: Biol Sci. 2020;375:20190341.
17. Guio L, Viei a C, González J. S ess a ec s he epigene ic ma ks added by
na u al ansposable elemen inse ions in D osophila melanogas e . Sci Rep.
2018;8:1–10.
18. Casacube a E, González J. The impac o ansposable elemen s in en i on-
men al adap a ion. Mol Ecol. 2013;22:1503–17.
19. Se a o-Capuchina A, Ma u e DR. The ole o ansposable elemen s in specia-
ion. Genes (Basel). 2018;9:254.
20. Co onado-Zamo a M, González J. T ansposons con ibu e o he unc ional
di e si ica ion o he head, gu , and o a y ansc ip omes ac oss D osophila
na u al s ains. Genome Res. 2023;33:1541–53.
21. Me enciano M, Ullas es A, de Ca a MAR, Ba on MG, Gonzalez J. Mul iple
independen e oelemen inse ions in he p omo e o a s ess esponse
gene ha e a iable molecula and unc ional e ec s in D osophila. PLoS
Gene . 2016;12:e1006249.
22. Me enciano M, González J. The in e play be ween de elopmen al s age and
en i onmen unde lies he adap i e e ec o a na u al ansposable elemen
inse ion. Mol Biol E ol. 2023;40:msad044.
23. Joly-Lopez Z, Bu eau TE. Exap a ion o ansposable elemen coding
sequences. Cu Opin Gene De . 2018;49:34–42.
24. Reis M, Viei a CP, La a R, Posnien N, Viei a J. O igin and consequences o
ch omosomal in e sions in he i ilis g oup o D osophila. Genome Biol E ol.
2018;10:3152–66.
25. Rius N, Delp a A, Ruiz A. A di e gen P elemen and i s associa ed MITE, BuT5,
gene a e ch omosomal in e sions and a e widesp ead wi hin he D osophila
eple a species g oup. Genome Biol E ol. 2013;5:1127–41.
26. Delp a A, Neg e B, Puig M, Ruiz A. The ansposon Galileo gene a es na u al
ch omosomal in e sions in D osophila by ec opic ecombina ion. PLoS ONE.
2009;4:e7883.
27. Ho mann AA, Riesebe g LH. Re isi ing he impac o in e sions in e olu ion:
om popula ion gene ic ma ke s o d i e s o adap i e shi s and specia ion?
Annu Re Ecol E ol Sys . 2008;39:21–42.
28. Jackson BC. Recombina ion-supp ession: how many mechanisms o ch o-
mosomal specia ion? Gene ica. 2011;139:393–402.
29. Fa ia R, Na a o A. Ch omosomal specia ion e isi ed: ea anging heo y wi h
pieces o e idence. T ends Ecol E ol. 2010;25:660–9.
30. Ayala D, Fon aine MC, Cohue A, Fon enille D, Vi alis R, Sima d F. Ch omo-
somal in e sions, na u al selec ion and adap a ion in he mala ia ec o
Anopheles Funes us. Mol Biol E ol. 2011;28:745–58.
31. E gen’e MB, Zelen so a H, Poluec o a H, Lyozin GT, Veleikod o skaja V,
Pya ko KI e al. Mobile elemen s and ch omosomal e olu ion in he i ilis
g oup o D osophila. P oceedings o he Na ional Academy o Sciences.
2000;97:11337–42.
32. Th ockmo on LH. The i ilis species g oup. Gene Bioogy D osophila.
1982;3:227–96.
33. MacMillan HA, Ande sen JL, Da ies SA, O e gaa d J. The capaci y o main ain
ion and wa e homeos asis unde lies in e speci ic a ia ion in D osophila
cold ole ance. Sci Rep. 2015;5:1–11.
34. Poikela N, Tyukmae a V, Hoikkala A, Kanka e M. Mul iple pa hs o cold ole -
ance: he ole o en i onmen al cues, mo phological ai s and he ci cadian
clock gene ille. BMC Ecol E ol. 2021;21:1–20.
35. Hoikkala A, Poikela N. Adap a ion and ecological specia ion in seasonally
a ying en i onmen s a high la i udes: D osophila i ilis g oup. Fly (Aus in).
Taylo and F ancis L d.; 2022. pp. 85–104.
36. Mi ol PM, Schä e MA, O sini L, Rou u J, Schlö e e C, Hoikkala A, e al. Phylo-
geog aphic pa e ns in D osophila mon ana. Mol Ecol. 2007;16:1085–97.
37. Kanka e M, Salminen T, Laiho A, Vesala L, Hoikkala A. Changes in gene exp es-
sion linked wi h adul ep oduc i e diapause in a no he n mal ly species: a
candida e gene mic oa ay s udy. BMC Ecol. 2010;10:1–9.
38. Kanka e M, Pa ke DJ, Me isalo M, Salminen TS. T ansc ip ional di e ences
be ween Diapausing and Non-diapausing D. Mon ana emales ea ed unde
he same pho ope iod and empe a u e. 2016;1–18.
39. Kau anen H, Kinnunen J, Hiillos A-L, Lankinen P, Hopkins D, Wibe g RAW, e al.
Selec ion o ep oduc ion unde sho pho ope iods changes diapause-
associa ed ai s and induces widesp ead genomic di e gence. J Exp Biol.
2019;222:jeb205831.
40. Pa ke DJ, Ri chie MG, Kanka e M. P epa ing o win e : he ansc ip omic
esponse associa ed wi h di e en day leng hs in D osophila mon ana. G3
Genes|Genomes|Gene ics. 2016;6:1373–81.
41. Pa ke DJ, Vesala L, Ri chie MG, Laiho A, Hoikkala A, Kanka e M. How consis-
en a e he ansc ip ome changes associa ed wi h cold acclima ion in wo
species o he D osophila i ilis g oup? He edi y (Edinb). 2015;115:13–21.
42. Pa ke DJ, Wibe g RAW, T i edi U, Tyukmae a VI, Gha bi K, Bu lin RK, e al.
In e and in aspeci ic genomic di e gence in D osophila mon ana shows
e idence o cold adap a ion. Genome Biol E ol. 2018;10:2086–101.
43. Wibe g RAW, Tyukmae a V, Hoikkala A, Ri chie MG, Kanka e M. Cold adap a-
ion d i es popula ion genomic di e gence in he ecological specialis ,
D osophila mon ana. Mol Ecol. 2021;30:3783–96.
Page 20 o 21Tahami e al. Mobile DNA (2024) 15:18
44. Mo ales-Hojas R, Päällysaho S. Compa a i e poly ene ch omosome maps o
D. mon ana. Ch omosoma. 2007;116(1):21–7.
45. S one WS, Gues WC, Wilson FD. The e olu iona y implica ions o he cy o-
logical polymo phism and phylogeny o he i ilis g oup o D osophila. P oc
Na l Acad Sci. 1960;46:350–61.
46. Poikela N, Lae sch DR, Hoikkala V, Lohse K, Kanka e M. Ch omosomal
In e sions and he Demog aphy o Specia ion in D osophila mon ana and
D osophila la omon ana. Gonzalez J, edi o . Genome Biol E ol. 2024;3:16.
47. Ruan J, Li H. Fas and accu a e long- ead assembly wi h w dbg2. Na Me h-
ods. 2020;17:155–8.
48. Zimin AV, Puiu D, Luo M-C, Zhu T, Ko en S, Ma çais G, e al. Hyb id assembly
o he la ge and highly epe i i e genome o Aegilops auschii, a p ogeni o
o b ead whea , wi h he MaSuRCA mega- eads algo i hm. Genome Res.
2017;27:787–92.
49. Chak abo y M, Baldwin-B own JG, Long AD, Eme son JJ. Con iguous and
accu a e de no o assembly o me azoan genomes wi h modes long ead
co e age. Nucleic Acids Res. 2016;44:e147–147.
50. Walke BJ, Abeel T, Shea T, P ies M, Abouelliel A, Sak hikuma S, e al. Pilon: an
in eg a ed ool o comp ehensi e mic obial a ian de ec ion and genome
assembly imp o emen . PLoS ONE. 2014;9:e112963.
51. Guan D, McCa hy SA, Wood J, Howe K, Wang Y, Du bin R. Iden i ying and
emo ing haplo ypic duplica ion in p ima y genome assemblies. Bioin o -
ma ics. 2020;36:2896–8.
52. Lae sch DR, Blax e ML, BlobTools. In e oga ion o genome assemblies.
F1000Res. 2017;6:1287.
53. Seppey M, Manni M, Zdobno EM. BUSCO: assessing genome assembly and
anno a ion comple eness. Gene P edic ion: Me hods P o ocols. 2019;227–45.
54. Alonge M, Lebeigle L, Ki sche M, Jenike K, Ou S, Aganezo S, e al. Au oma ed
assembly sca olding using RagTag ele a es a new oma o sys em o high-
h oughpu genome edi ing. Genome Biol. 2022;23:1–19.
55. Li H. Minimap2: pai wise alignmen o nucleo ide sequences. Bioin o ma ics.
2018;34:3094–100.
56. Flu e T, Dup a E, Feuille C, Quesne ille H. Conside ing ansposable
elemen di e si ica ion in de no o anno a ion app oaches. PLoS ONE.
2011;6:e16526.
57. Bao Z, Eddy SR. Au oma ed de no o iden i ica ion o epea sequence ami-
lies in sequenced genomes. Genome Res. 2002;12:1269–76.
58. Malik L, Almoda esi F, Pa o R. G oupe : g aph-based clus e ing and
anno a ion o imp o ed de no o ansc ip ome analysis. Bioin o ma ics.
2018;34:3265–72.
59. Edga RC, Mye s EW. PILER: iden i ica ion and classi ica ion o genomic
epea s. Bioin o ma ics-Ox o d. 2005;21:i152.
60. Ju ka J, Kapi ono VV, Pa licek A, Klonowski P, Kohany O, Walichiewicz J.
Repbase Upda e, a da abase o euka yo ic epe i i e elemen s. Cy ogene
Genome Res. 2005;110:462–7.
61. Tho aldsdó i H, Robinson JT, Mesi o JP. In eg a i e Genomics Viewe
(IGV): high-pe o mance genomics da a isualiza ion and explo a ion. B ie
Bioin o m. 2013;14:178–92.
62. Wicke T, Sabo F, Hua-Van A, Benne zen JL, Capy P, Chalhoub B, e al. A
uni ied classi ica ion sys em o euka yo ic ansposable elemen s. Na Re
Gene . 2007;8:973–82.
63. O ozco S, Sie a P, Du bin R, Gonzalez J. MCHelpe au oma ically cu a es
ansposable elemen lib a ies ac oss species. bioRxi . 2023;2010–23.
64. Smi AFA, Hubley R, G een P. Repea Maske Open-4.0. 2013.
65. Pa ke DJ, En all T, Ri chie MG, Kanka e M. Sex-speci ic esponses o cold
in a e y cold- ole an , no he n D osophila species. He edi y (Edinb).
2021;126:695–705.
66. Chen S, Zhou Y, Chen Y, Gu J. Fas p: an ul a- as all-in-one FASTQ p ep oces-
so . Bioin o ma ics. 2018;34:i884–90.
67. Dobin A, Da is CA, Schlesinge F, D enkow J, Zaleski C, Jha S, e al. STAR:
ul a as uni e sal RNA-seq aligne . Bioin o ma ics. 2013;29:15–21.
68. B ůna T, Ho KJ, Lomsadze A, S anke M, Bo odo sky M. BRAKER2: au oma ic
euka yo ic genome anno a ion wi h GeneMa k-EP + and AUGUSTUS sup-
po ed by a p o ein da abase. NAR Genom Bioin o m. 2021;3:lqaa108.
69. S anke M, Schö mann O, Mo gens e n B, Waack S. Gene p edic ion in euka y-
o es wi h a gene alized hidden Ma ko model ha uses hin s om ex e nal
sou ces. BMC Bioin o ma ics. 2006;7:1–11.
70. S anke M, Diekhans M, Bae sch R, Haussle D. Using na i e and syn enically
mapped cDNA alignmen s o imp o e de no o gene inding. Bioin o ma ics.
2008;24:637–44.
71. Ba ne DW, Ga ison EK, Quinlan AR, S ömbe g MP, Ma h GT. BamTools: a
C + + API and oolki o analyzing and managing BAM iles. Bioin o ma ics.
2011;27:1691–2.
72. Lomsadze A, Bu ns PD, Bo odo sky M. In eg a ion o mapped RNA-Seq eads
in o au oma ic aining o euka yo ic gene inding algo i hm. Nucleic Acids
Res. 2014;42:e119–119.
73. Buch ink B, Xie C, Huson DH. Fas and sensi i e p o ein alignmen using
DIAMOND. Na Me hods. 2015;12:59–60.
74. Ho KJ, Lange S, Lomsadze A, Bo odo sky M, S anke M. BRAKER1: unsupe -
ised RNA-Seq-based genome anno a ion wi h GeneMa k-ET and AUGUS-
TUS. Bioin o ma ics. 2016;32:767–9.
75. Shuma e A, Salzbe g SL. Li o : accu a e mapping o gene anno a ions.
Bioin o ma ics. 2021;37:1639–43.
76. Li H. Aligning sequence eads, clone sequences and assembly con igs wi h
BWA-MEM. a Xi p ep in a Xi :13033997. 2013.
77. Ta aso A, Vilella AJ, Cuppen E, Nijman IJ, P ins P. Sambamba: as p ocessing
o NGS alignmen o ma s. Bioin o ma ics. 2015;31:2032–4.
78. Ga ison E, Ma h G. Haplo ype-based a ian de ec ion om sho - ead
sequencing. a Xi P ep in a Xi :12073907. 2012.
79. Tan A, Abecasis GR, Kang HM. Uni ied ep esen a ion o gene ic a ian s.
Bioin o ma ics. 2015;31:2202–4.
80. Pu cell S, Neale B, Todd-B own K, Thomas L, Fe ei a MAR, Bende D, e al.
PLINK: a ool se o whole-genome associa ion and popula ion-based link-
age analyses. Am J Hum Gene . 2007;81:559–75.
81. Quinlan AR, Hall IM. BEDTools: a lexible sui e o u ili ies o compa ing
genomic ea u es. Bioin o ma ics. 2010;26:841–2.
82. Wickham H, Chang W, Wickham MH. Package ‘ggplo 2.’ C ea e elegan da a
isualisa ions using he g amma o g aphics Ve sion. 2016;2:1–189.
83. Va gas-Cha ez C, Pendy NML, Nsango SE, Aguile a L, Ayala D, González
J. T ansposable elemen a ian s and hei po en ial adap i e impac in
u ban popula ions o he mala ia ec o Anopheles coluzzii. Genome Res.
2022;32:189–202.
84. Ko le R, Gómez-Sánchez D, Schlö e e C. PoPoola ionTE2: compa a i e
popula ion genomics o ansposable elemen s using Pool-Seq. Mol Biol E ol.
2016;33:2759–64.
85. Yu T, Huang X, Dou S, Tang X, Luo S, Theu kau WE, e al. A benchma k and an
algo i hm o de ec ing ge mline ansposon inse ions and measu ing de
no o ansposon inse ion equencies. Nucleic Acids Res. 2021;49:e44–44.
86. G ama es LS, Agapi e J, A ill H, Cal i BR, C osby MA, Dos San os G, e al.
FlyBase: a guided ou o highligh ed ea u es. Gene ics. 2022;220:iyac035.
87. Villanue a-Cañas JL, Ho a h V, Aguile a L, González J. Di e se amilies o
ansposable elemen s a ec he ansc ip ional egula ion o s ess- esponse
genes in D osophila melanogas e . Nucleic Acids Res. 2019;47:6842–57.
88. FlyBase. Da abase [In e ne ]. [ci ed 2022 No 20]. A ailable om: h ps:// ly-
base.o g.
89. Salminen TS, Vesala L, Laiho A, Me isalo M, Hoikkala A, Kanka e M. Seasonal
gene exp ession kine ics be ween diapause phases in D osophila i ilis
g oup species and o e win e ing di e ences be ween diapausing and non-
diapausing emales. Sci Rep. 2015;5:1–13.
90. Guo Z, Qin J, Zhou X, Zhang Y. Insec ansc ip ion ac o s: a landscape o
hei s uc u es and biological unc ions in D osophila and beyond. In J Mol
Sci. 2018;19:3691.
91. Jukam D, Vie s K, Ande son C, Zhou C, DeFo d P, Yan J, e al. The insula o
p o ein BEAF-32 is equi ed o Hippo pa hway ac i i y in he e minal di -
e en ia ion o neu onal sub ypes. De elopmen . 2016;143:2389–97.
92. Reddy DA, P asad B, Mi a CK. Func ional classi ica ion o ansc ip ion
ac o binding si es: in o ma ion con en as a me ic. J In eg Bioin (JIB).
2006;3:32–44.
93. Sandelin A, Alkema W, Engs öm P, Wasse man WW, Lenha d B. JASPAR: an
open-access da abase o euka yo ic ansc ip ion ac o binding p o iles.
Nucleic Acids Res. 2004;32:D91–4.
94. G an CE, Bailey TL, Noble WS. FIMO: scanning o occu ences o a gi en
mo i . Bioin o ma ics. 2011;27:1017–8.
95. Bailey TL, Boden M, Buske FA, F i h M, G an CE, Clemen i L, e al. MEME SUITE:
ools o mo i disco e y and sea ching. Nucleic Acids Res. 2009;37:W202–8.
96. Bailey TL, Johnson J, G an CE, Noble WS. The MEME sui e. Nucleic Acids Res.
2015;43:W39–49.
97. Slou skin A, Danino YM, O ens ein Y, Zeha i Y, Donige T, Shami R, e al. Ele-
meNT: a compu a ional ool o de ec ing co e p omo e elemen s. T ansc ip-
ion. 2015;6:41–50.
98. ElemeNT co e p omo e de ec ing ool [In e ne ]. [ci ed 2022 Sep 11]. h p://
li e acul y.biu.ac.il/ge shon- ama /index.php/ esou ces
Page 21 o 21Tahami e al. Mobile DNA (2024) 15:18
99. Hackl T, Ankenb and M, an Ad ichem B. gggenomes: A G amma o G aph-
ics o Compa a i e Genomics [In e ne ]. 2024 [ci ed 2024 Jun 20]. h ps://
gi hub.com/ hackl/gggenomes
100. Li H, Du bin R. Fas and accu a e sho ead alignmen wi h Bu ows–Wheele
ans o m. Bioin o ma ics. 2010;25:1754–60.
101. Neph S, Kuehn MS, Reynolds AP, Haugen E, Thu man RE, Johnson AK, e al.
BEDOPS: high-pe o mance genomic ea u e ope a ions. Bioin o ma ics.
2012;28:1919–20.
102. Pica d pipeline. B oad Ins i u e [In e ne ]. [ci ed 2023 Aug 27]. h p://b oadin-
s i u e.gi hub.io/pica d
103. Sedlazeck FJ, Reschenede P, Smolka M, Fang H, Na es ad M, Von Haesele A,
e al. Accu a e de ec ion o complex s uc u al a ia ions using single-mole-
cule sequencing. Na Me hods. 2018;15:461–8.
104. Rausch T, Zichne T, Schla l A, S ü z AM, Benes V, Ko bel JO. DELLY: s uc u al
a ian disco e y by in eg a ed pai ed-end and spli - ead analysis. Bioin o -
ma ics. 2012;28:i333–9.
105. Je a es DC, Jolly C, Ho i M, Speed D, Shaw L, Rallis C, e al. T ansien s uc u al
a ia ions ha e s ong e ec s on quan i a i e ai s and ep oduc i e isola-
ion in ission yeas . Na Commun. 2017;8:14061.
106. Rech GE, Radío S, Gui ao-Rico S, Aguile a L, Ho a h V, G een L, e al. Popula-
ion-scale long- ead sequencing unco e s ansposable elemen s associa ed
wi h gene exp ession a ia ion and adap i e signa u es in D osophila. Na
Commun. 2022;13:1948.
107. Fonseca PM, Mou a RD, Wallau GL, Lo e o ELS. The mobilome o D osophila
incomp a, a lowe -b eeding species: compa ison o ansposable ele-
men landscapes among gene alis and specialis lies. Ch omosome Res.
2019;27:203–19.
108. Sim C, Denlinge DL. Insulin signaling and FOXO egula e he o e -
win e ing diapause o he mosqui o Culex pipiens. P oc Na l Acad Sci.
2008;105:6777–81.
109. Dan o W, Da is MM, Lind all JM, Tang X, U ell H, Junell A, e al. The Oc 1
homolog nubbin is a ep esso o NF-κB-dependen immune gene exp es-
sion ha inc eases he ole ance o gu mic obio a. BMC Biol. 2013;11:1–18.
110. Roseman RR, Pi o a V, Geye PK. The su (hw) p o ein insula es exp ession
o he D osophila melanogas e whi e gene om ch omosomal posi ion-
e ec s. EMBO J. 1993;12:435–42.
111. Tapanainen R, Pa ke DJ, Kanka e M. Pho osensi i e Al e na i e Splicing o
he ci cadian clock gene imeless is Popula ion Speci ic in a Cold-adap ed ly,
D osophila mon ana. G3 Genes|Genomes|Gene ics. 2018;8:1291–7.
112. Obba d DJ, Maclennan J, Kim K-W, Rambau A, O’G ady PM, Jiggins FM. Es i-
ma ing di e gence da es and subs i u ion a es in he D osophila phylogeny.
Mol Biol E ol. 2012;29:3459–73.
113. Yusu LH, Tyukmae a V, Hoikkala A, Ri chie MG. Di e gence and in og ession
among he i ilis g oup o D osophila. E ol Le . 2022;6:537–51.
114. Lo e o ELS, Ca a e o CMA, Capy P. Re isi ing ho izon al ans e o anspos-
able elemen s in D osophila. He edi y (Edinb). 2008;100:545–54.
115. McCulle s TJ, S einige M. T ansposable elemen s in D osophila. Mob Gene
Elem. 2017;7:1–18.
116. Rius N, Guillén Y, Delp a A, Kapus a A, Fescho e C, Ruiz A. Explo a ion o he
D osophila buzza ii ansposable elemen con en sugges s unde es ima ion
o epea s in D osophila genomes. BMC Genomics. 2016;17:1–14.
117. Maumus F, Fis on-La ie A-S, Quesne ille H. Impac o ansposable elemen s
on insec genomes and biology. Cu Opin Insec Sci. 2015;30–6.
118. Ellison CE, Bach og D. Dosage compensa ion ia ansposable elemen
media ed ewi ing o a egula o y ne wo k. Science (1979). 2013;342:846–50.
119. Viei a C, Na don C, A pin C, Lepe i D, Biémon C. E olu ion o genome size
in D osophila. Is he in ade ’s genome being in aded by ansposable ele-
men s? Mol Biol E ol. 2002;19:1154–61.
120. S i C, Go don SP, Wicke T, Vogel JP, Roulin AC. Recen ac i i y in expanding
popula ions and pu i ying selec ion ha e shaped ansposable elemen land-
scapes ac oss na u al accessions o he Medi e anean g ass B achypodium
dis achyon. Genome Biol E ol. 2018;10:304–18.
121. Me enciano M, Iacome i C, González J. A unique clus e o oo inse ions in
he p omo e egion o a s ess esponse gene in D osophila melanogas e .
Mob DNA. 2019;10:1–11.
122. Psenako a K, Kohou o a K, Obsilo a V, Ausse lechne MJ, Ve e ka V, Obsil T.
Fo khead domains o FOXO ansc ip ion ac o s di e in bo h o e all con o -
ma ion and dynamics. Cells. 2019;8:966.
123. Kapun M, Fla T. The adap i e signi icance o ch omosomal in e sion poly-
mo phisms in D osophila melanogas e . Mol Ecol. 2019;28:1263–82.
124. Poikela N, Lae sch DR, Kanka e M, Hoikkala A, Lohse K. Expe imen al in o-
g ession in D osophila: asymme ic pos zygo ic isola ion associa ed wi h
ch omosomal in e sions and an incompa ibili y locus on he X ch omosome.
Mol Ecol. 2023;32:854–66.
Publishe ’s no e
Sp inge Na u e emains neu al wi h ega d o ju isdic ional claims in
published maps and ins i u ional a ilia ions.