scieee Open visual document viewer

Transposable elements in Drosophila montana from harsh cold environments

Tahami, Mohadeseh S.,Vargas-Chavez, Carlos,Poikela, Noora,Coronado-Zamora, Marta,González, Josefa,Kankare, Maaria

Full text

This is a sel -a chi ed e sion o an o iginal a icle. This e sion may di e om he o iginal in pagina ion and ypog aphic de ails. Au ho (s): Ti le: Yea : Ve sion: Copy igh : Righ s: Righ s u l: Please ci e he o iginal e sion: CC BY-NC-ND 4.0 h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ T ansposable elemen s in D osophila mon ana om ha sh cold en i onmen s © The Au ho (s) 2024 Published e sion Tahami, Mohadeseh S.; Va gas-Cha ez, Ca los; Poikela, Noo a; Co onado- Zamo a, Ma a; González, Jose a; Kanka e, Maa ia Tahami, M. S., Va gas-Cha ez, C., Poikela, N., Co onado-Zamo a, M., González, J., & Kanka e, M. (2024). T ansposable elemen s in D osophila mon ana om ha sh cold en i onmen s. Mobile DNA, 15, A icle 18. h ps://doi.o g/10.1186/s13100-024-00328-7 2024 RESEARCH Open Access © The Au ho (s) 2024. Open Access This a icle is licensed unde a C ea i e Commons A ibu ion-NonComme cial-NoDe i a i es 4.0 In e na ional License, which pe mi s any non-comme cial use, sha ing, dis ibu ion and ep oduc ion in any medium o o ma , as long as you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons licence, and indica e i you modi ied he licensed ma e ial. You do no ha e pe mission unde his licence o sha e adap ed ma e ial de i ed om his a icle o pa s o i . The images o o he hi d pa y ma e ial in his a icle a e included in he a icle’s C ea i e Commons licence, unless indica ed o he wise in a c edi line o he ma e ial. I ma e ial is no included in he a icle’s C ea i e Commons licence and you in ended use is no pe mi ed by s a u o y egula ion o exceeds he pe mi ed use, you will need o ob ain pe mission di ec ly om he copy igh holde . To iew a copy o his licence, isi h p:// c ea i ecommons.o g/licenses/by-nc-nd/4.0/. Tahami e al. Mobile DNA (2024) 15:18 h ps://doi.o g/10.1186/s13100-024-00328-7 Mobile DNA *Co espondence: Jose a González [email p o ec ed] Maa ia Kanka e maa ia.kanka [email p o ec ed] 1Depa men o Biological and En i onmen al Science, Uni e si y o Jy äskylä, Jy äskylä, Finland 2Ins i u e o E olu iona y Biology, CSIC, UPF, Ba celona, Spain 3Cen e o Biological Di e si y, Uni e si y o S And ews, S And ews, UK 4Ins i u Bo ànic de Ba celona (IBB), CSIC-CMCNB, Ba celona 08038, Ca alonia, Spain Abs ac Backg ound Subs an ial disco e ies du ing he pas cen u y ha e e ealed ha ansposable elemen s (TEs) can play a c ucial ole in genome e olu ion by a ec ing gene exp ession and inducing gene ic ea angemen s, among o he molecula and s uc u al e ec s. Ye , ou knowledge on he ole o TEs in adap a ion o ex eme clima es is s ill a i s in ancy. The a ailabili y o long- ead sequencing has opened up he possibili y o iden i y and s udy po en ial unc ional e ec s o TEs wi h highe p ecision. In his wo k, we used D osophila mon ana as a model o cold-adap ed o ganisms o s udy he associa ion be ween TEs and adap a ion o ha sh clima es. Resul s Using he PacBio long- ead sequencing echnique, we de no o iden i ied and manually cu a ed TE sequences in i e D osophila mon ana genomes om eco-geog aphically dis inc popula ions. We iden i ied 489 new TE consensus sequences which ep esen ed 92% o he o al TE consensus in D. mon ana. O e all, 11–13% o he D. mon ana genome is occupied by TEs, which as expec ed a e non- andomly dis ibu ed ac oss he genome. We iden i ied i e po en ially ac i e TE amilies, mos o hem om he e o ansposon class o TEs. Addi ionally, we ound TEs p esen in he i e analyzed genomes ha we e loca ed nea by p e iously iden i ied cold ole an genes. Some o hese TEs con ain p omo e elemen s and ansc ip ion binding si es. Finally, we de ec ed TEs nea by ixed and polymo phic in e sion b eakpoin s. Conclusions Ou esea ch e ealed a signi ican numbe o newly iden i ied TE consensus sequences in he genome o D. mon ana, sugges ing ha non-model species should be s udied o ge a comp ehensi e iew o he TE epe oi e in D osophila species and beyond. Genome anno a ions wi h he new D. mon ana lib a y allowed us o iden i y TEs loca ed nea by cold ole an genes, and p esen a high popula ion equencies, ha con ain egula o y egions and a e hus good candida es o play a ole in D. mon ana cold s ess esponse. Finally, ou anno a ions also allow us o iden i y o he i s ime TEs p esen in he b eakpoin s o h ee D. mon ana in e sions. Keywo ds D osophila mon ana, Cold adap a ion, T ansposable elemen s, Ac i e TEs, Ch omosomal in e sions T ansposable elemen s in D osophila mon ana om ha sh cold en i onmen s Mohadeseh S.Tahami1, Ca losVa gas-Cha ez2, Noo aPoikela1,3, Ma aCo onado-Zamo a2,4, Jose aGonzález2,4* and Maa iaKanka e1* Page 2 o 21Tahami e al. Mobile DNA (2024) 15:18 Backg ound En i onmen al s ess is one o he key ai s a ec ing species su i al and dis ibu ion, especially in no he n a eas whe e empe a u e luc ua ions a e less p edic - able and ex eme wea he e en s a e becoming mo e equen because o clima e change [1, 2]. Recen s ud- ies ha e indica ed ha species may adap h ough a i- ous in e plays be ween di e en gene ic elemen s in he o ganism’s genome [3–7]. One o such elemen s d i - ing adap a ion a e ansposable elemen s (TEs). TEs a e epe i i e sequences wi h he abili y o eplica e hem- sel es independen ly o he genome and change hei posi ion he eby gene a ing di e se mu a ions. They ep- esen a signi ican po ion o he genome in many spe- cies, o example, o e wo hi ds o he human genome [8] and 90% o he whea genome [9] a e TEs. In insec s, TE con en also con ibu es signi ican ly o genome size a ia ion [10], o example, in D osophila lies, TE con- en can a y be ween 3% and 30% [11]. The ac i i y and he abundance o hese highly epe i i e sequences in he genome sugges s ha TEs could be impo an ac o s in he e olu ion o many species. TEs can impac species adap a ion o en i onmen al changes. Fo example, hey a e known o be ac i a ed unde s ess [12] and can a ec he species esponse o en i onmen al s esso s such as hea shock [13], cold [14], pes icides [15], insec icides [16] and oxida i e s ess [17] hus c ea ing a ole ance agains hese s esso s. TEs can ha e di e se e ec s on he hos ’s genome based on he loca ion hey a e inse ed and he egula o y sequences hey con ain, e.g. hey can c ea e inse ional mu a ions, up- o down egula e nea by genes, o c ea e new gene a ian s by in oducing new exons, new s op codons, o al e na i e splice si es [18–20]. In some cases, one o hese newly gene a ed alleles may become adap- i e in esponse o en i onmen al s ess. This so-called ‘adap i e TE’ will be co-op ed and hus ole a ed by na u- al selec ion. The case o he co-op ed FB i0019985 oo solo-LTR ansposon, loca ed in he p omo e o he Lime gene o D. melanogas e is an example in which he adap i e TE a ec s gene exp ession associa ed wi h cold-s ess and immune ole ance [21, 22]. In he case o immune ole ance, he au ho s show ha FB i0019985 oo solo-LTR inse ion is adding unc ional ansc ip- ion ac o binding si es (TFBS) o ansc ip ion ac o s in ol ed in immune esponse [22]. I he s ess condi- ions pe sis , co-op ed TEs migh become ixed in he popula ion [23]. Ano he way ha TEs can pa icipa e in en i onmen al adap a ion is by inducing ec opic ecombina ion leading o s uc u al a ia ions such as in e sions [24, 25]. One such case is epo ed o a Galileo ansposon ha has gene a ed h ee polymo phic in e sions in D osophila buzza ii h ough ec opic ecombina ion [26]. In e sions in u n can play a ole in species adap a ion and specia- ion in many ways [ e iewed in e.g. 27–29]. Thei main e olu iona y signi icance is ha hey can educe ecom- bina ion be ween a o able combina ions o alleles while p o ec ing se s o locally adap ed genes, he eby p omo - ing ecological di e gence, and os e ing ep oduc i e isola ion wi hin species [30]. Fo example, in e sions a e p e ailing d i e s o popula ion di e gence in D. i ilis species g oup [31]. Despi e he abundance o bo h in a- speci ic polymo phic and in e speci ic ixed in e sions in Dip e a, s udies on he ole o TEs on ch omosomal in e sions a e s ill sca ce. The no he n mal ly, D osophila mon ana, is a widely dis ibu ed insec species which belongs o he D osoph- ila i ilis g oup [32] and is one o he mos cold- ole an D osophila species [33, 34]. These no he n mal lies can spend up o 6 mon hs in subze o empe a u es and se - e al popula ions ha e adap ed o li e in la i udes abo e he A c ic Ci cle in No he n Scandina ia. In lowe la i- udes, lies inhabi high ele a ions (abou 3,000m) e.g., in he Rocky Moun ains in Colo ado (USA). They also exis in he wa me coas al a eas in Washing on and O e- gon (USA) and hence show a wide ange o eco-clima e adap a ion [35]. Popula ions o D. mon ana lies a e clus- e ed in o Eu opean, No h Ame ican and Asian popula- ions [36] and he cu en geog aphical pa e n is likely he esul o sp eading No h Ame ican popula ions in o Eu asia h ough he Be ing S ai be ween 450,000 and 1,750,000 yea s ago [35] esul ing in i s unique in aspe- ci ic di e si ica ion. In D. mon ana, adap a ion o he ex eme clima e is modula ed by se e al genes associa ed wi h cold ole ance [37–43]. These genes egula e a a i- e y o unc ions om gene al cold esis ance o cu icu- la and ol ac o y p ocesses and pho ope iodic diapause, enhancing he species o e win e ing abili y [34, 42]. I is also possible ha some TEs may ha e a ole in he egula- ion o he cold- ela ed genes, leading o he species high esis ance, e.g. o eezing empe a u es. Mo eo e , based on ea lie poly ene ch omosome s udies, D. mon ana ca ies se e al ixed and polymo phic in e sions [44, 45]. So a , one ixed in e sion [46] (in compa ison o D o- sophila la omon ana), and wo polymo phic in e sions (Poikela e al., in p ep.) has been cha ac e ized a he genomic le el, bu he link be ween TEs and in e sions in D. mon ana has no been s udied ye . He e, we iden i ied TEs in D. mon ana lies ac oss i s dis ibu ion ange o he i s ime using long- ead genome sequence da a. To s udy he po en ial ole o TEs in adap a ion o no he n en i onmen s, we i s p o- duced a comp ehensi e manually cu a ed TE lib a y om No h Ame ican, No h Eu opean and Asian popula ions o D. mon ana. Nex , we in es iga ed he abundance, densi y, dis ibu ion and ac i i y o TEs h oughou hese genomes. Finally, we s udied he associa ion be ween TEs Page 3 o 21Tahami e al. Mobile DNA (2024) 15:18 and ch omosomal in e sions, as well as he link be ween TEs and a selec ed se o cold ole an candida e genes ound in p e ious s udies o iden i y possible adap i e TEs. We add essed he ollowing ques ions: (i) wha is he genome-wide TE p o ile in D. mon ana and does i di e om ha o o he D osophila species? (ii) Can we de ec po en ially ac i e TEs in he D. mon ana’s genome? (iii) Can we ind any e idence o he ole o TEs in adap a- ion o he no he n en i onmen s? and (i ) a e he e TE inse ions loca ed nea by in e sion b eakpoin s? Me hods Sample collec ions D osophila mon ana lies we e collec ed om 13 loca- ions in No h Ame ica (NA), No h Eu ope (NE), and Fa Eas Asia (FE) be ween 2013 and 2021 (Fig.1 and Supplemen a y Table S1). The collec ed emales we e b ough o he ly labo a o y o he Uni e si y o Jy äs- kylä, Finland, and kep in cons an ligh , 19 ˚C and ~ 60% humidi y. The emales ha had ma ed in na u e we e allowed o lay eggs in mal ials o se e al days. The eme ged F1 p ogeny o each emale we e kep oge he o p oduce he nex gene a ion and o es ablish iso emale s ains. A e he es ablishmen o iso emale s ains, all wild-collec ed emales and hei F1 p ogeny, excep hose collec ed in FE, we e s o ed in 70% E OH a -20 ˚C. Fo PacBio (Paci ic Biosciences) long- ead sequenc- ing, D. mon ana emales we e collec ed om i e iso e- male s ains o igina ing om i e loca ions: h ee in NA, Sewa d (monSE13F37; Alaska, USA) and Jackson (mon- JX13F48; Wyoming, USA) om a p e ious s udy [46] and C es ed Bu e (mon34CC5; Colo ado, USA) om he cu en s udy; one in NE, Oulanka (monOU13F149; Fin- land); and one in FE, Kamcha ka (monKR1309; Russia) (Fig.1 and Supplemen a y Tables S1 & S2). P io o col- lec ing he lies o PacBio sequencing, iso emale s ains we e kep in he labo a o y o ~ 50 gene a ions (Supple- men a y Table S2). Fo Illumina sho - ead sequencing, we used he s o ed wild-collec ed emales o hei F1 daugh e s om 12 col- lec ion si es in NA and NE. Because he wild-collec ed emales o hei F1 daugh e s om FE we e no s o ed a -20 ˚C, we pe o med Illumina sequencing on emales collec ed om iso emale s ains kep in he labo a o y o ~ 60 gene a ions (Fig.1 and Supplemen a y Tables S1 & S3). Long- ead sequencing PacBio sequencing echnique was used o he i e D. mon ana iso emale s ains. Excep o Sewa d and Jack- son [46], all samples we e sequenced in his s udy. Fo all samples, DNA was ex ac ed om a pool o 60 whole bodies o emale lies. De ails o samples, DNA ex ac- ion, and sequencing me hods a e gi en in Supplemen- a y Table S2. Sho - ead sequencing Illumina whole-genome sequencing was pe o med o 3–11 wild-collec ed o F1 emales pe popula ion om NA and NE, and o 3 emales collec ed om FE iso e- male s ains (Fig.1 and Supplemen a y Table S3). The majo i y o he ly samples (94/101) we e sequenced o his s udy, bu se en samples we e ob ained om Poikela e al. [46] (see Supplemen a y Table S3). DNA ex ac ions we e ca ied ou o single emales using ce yl ime h- ylammonium b omide (CTAB) solu ion wi h RNAse ea men , Phenol-Chlo o o m-Isoamyl alcohol (25:24:1) and Chlo o o m-Isoamyl alcohol (24:1) washing s eps and e hanol p ecipi a ion a he Uni e si y o Jy äskylä, Finland. De ails o lib a y p epa a ion, sequencing ech- nologies, loca ions, and yea s a e gi en in Addi ional ile 1, Supplemen a y Table S3. We gene a ed on a e age 19-97X co e age pe sample (Supplemen a y Table S3). De no o genome assemblies and sca olding We ob ained de no o genome assemblies o D. mon- ana Sewa d and Jackson om Poikela e al. [46] and cons uc ed new de no o genome assemblies o D. mon- ana C es ed Bu e, Kamcha ka and Oulanka ollowing he same p o ocol. In b ie , we assembled he genomes using PacBio and he espec i e Illumina eads wi h w dbg2 pipeline 2.5 [47] and MaSuRCA hyb id assem- ble 3.3.9 [48], a e which he assembly con igui y was imp o ed using quickme ge 0.3 [49] and he assemblies Fig. 1 Sample loca ions o he D. mon ana popula ions analyzed in his s udy. The ed ci cles indica e he loca ion o he i e genomes ha we e used o gene a e he o iginal TE lib a y (Supplemen a y Table S2). The ed and blue ci cles show he loca ion o popula ions used in he Illumina sequencing (Supplemen a y Table S3) Page 4 o 21Tahami e al. Mobile DNA (2024) 15:18 we e polished wi h he same Illumina eads using Pilon 1.23 [50]. Finally, uncollapsed he e ozygous egions we e emo ed using pu ge_dups 1.0.1 [51] and genomic con aminan s we e emo ed using BlobTools 1.1 [52]. We es ima ed he comple eness o he assemblies using he BUSCO pipeline 5.1.2 based on he Dip e a da abase “dip e a_odb10” [53], which sea ches o he p esence o 3,285 conse ed single copy Dip e a o hologues. Ch omosome-le el genome assembly o D. mon ana Sewa d was also ob ained om Poikela e al. [46]. This assembly was used o sca old he o he genomes (Jack- son, C es ed Bu e, Kamcha ka and Oulanka) using a e e ence-genome-guided sca olding ool RagTag 2.1.0 [54] which o ien s and o de s he inpu con igs based on a e e ence using minimap2 [55]. Building a de no o TE lib a y The i e de no o assembled genomes we e used o de no o iden i ica ion o TEs using he TEdeno o pipeline in he REPET package 3.0 [56]. In sho , epea ed egions we e de ec ed by sel -alignmen o he genomic chunks, he TE candida es we e clus e ed using h ee me hods: Recon 1.08 [57], G oupe 2.27 [58] and Pile 1.0 [59]. Consensus sequences we e de i ed om mul iple align- men s in each clus e . The consensuses we e classi ied based on s uc u al ea u es o homology ma ch wi h he Repbase e e ence lib a y 20.05 [60] and hidden Ma ko model (HMM) p o iles. To anno a e consensus TEs, each TE lib a y was blas ed agains i s hos genome using h ee di e en ools Blas e 2.25, Repea Maske 4.0.6, and CENSOR 4.1, in eg a ed in o TEanno pipe- line [56]. To apply a s a is ical il e ing, a andomized genomic chunk alignmen s was also applied in o de o calcula e he high-sco ing segmen pai s (HSP). The nex s eps il e ed and combined HSPs by keeping only he ones ha ing a sco e highe han he HSP h eshold, emo ed spu ious HSPs, and calcula ed iden i y pe cen - ages. We also kep and me ged he sho simple epea (SSR) da a by applying s ep 4 and 5 o TEanno . Then, consensus sequences wi h less han one ull leng h copy, iden i ied by TEanno as sequences wi h 95% o simila i y ma ch o e he ull leng h o TE consensus, h oughou he genome we e il e ed ou and he TEanno p ocess was epea ed on he upda ed lib a y. Fo manual cu a- ion, he consensus sequences o each indi idual lib a y we e isualized in he In eg a i e Genomics Viewe (IGV) 2.8.0 [61] along wi h hei s uc u al ea u es, conse ed domains, nucleo ide, and amino acid homology ma ches. We cu a ed iden i ied TEs based on Wicke ’s ea u es [62] and we ollowed se e al s eps o educe he alse posi i e e o in ou TE lib a y and o emo e he edun- dancy. Fo mo e de ails on he manual cu a ion p ocess please e e o he Addi ional ile 1, Supplemen a y Table S4 and Addi ional ile 2. To alida e ou manual cu a ion p ocess, we un MCHelpe on he inal lib a y [63]. He e och oma in iden i ica ion Because we we e in e es ed in TE inse ions wi h po en- ial unc ional impac , we iden i ied he e och oma in egions o be emo ed om u he TE analyses. A e he manual TE cu a ion, he inal lib a y was used o anno a e TEs ac oss each ch omosome-le el assembly using Repea Maske 4.1.0 [64] wi h ollowing pa ame- e s: -e ncbi -nolow -xsmall -gccalc. Since no p io in o - ma ion on he he e och oma in egions o D. mon ana genome was a ailable, we plo ed he o al TE densi ies pe 5kb genomic bins. We conside ed he egions wi h he highes TE densi y o be he e och oma ic egions. To iew he TE densi y plo please e e o Addi ional ile 3 Supplemen a y Figu e S1 and o he able he e o- s. euch oma in coo dina es, e e o Addi ional ile 1, Supplemen a y Table S5. We conside ed a egion he e o- ch oma in i i was spanned by a leas 4 bins wi h o e 80% TE densi y. As expec ed, we ound a nega i e co - ela ion be ween TE con en and gene densi y in he e o- ch oma ic egions (Supplemen a y Table S6A). Gene anno a ion We used he Sewa d ch omosome-le el assembly as he e e ence o genome anno a ion. The cu a ed TE lib a y o D. mon ana was used o mask genomes using Repea Maske 4.1.0 wi h he ollowing pa ame e s: -e ncbi -nolow -xsmall -gccalc. The so -masked genome was anno a ed using D. mon ana RNAseq da a (Illu- mina T useq 150bp PE) [65]. RNA-seq eads we e i s immed using as p 0.20.0 [66] and mapped agains he so -masked genome using STAR 2.7.10a [67]. Gene anno a ions we e ca ied ou wi h b ake 2.1.6 [68] wi h de aul pa ame e s, wi h RNAseq as e idence using Augus us 3.4.0 and GeneMa k-ET 4.48 implemen ed pipelines [69–74]. Finally, he anno a ion was ans e ed o he o he ch omosome-le el genomes using Li o 1.6.3 [75]. A ange o 14,504 o 14,725 genes we e ans- e ed o each genome (Supplemen a y Table S6B). The a e age e o a e was calcula ed based on he o al num- be o exons ha we e co ec ly ans e ed o each gene o be a 1.1%. SNP calling To examine i SNP a ia ion wi hin TE egions can e lec he geog aphical a ia ion be ween D. mon ana genomes ac oss i s dis ibu ion, Illumina eads we e used o ex ac nucleo ide a ian s om he whole genome and om TE egions o compa ison. Reads we e il e ed o adap - e s and bases below quali y o 20 we e immed a bo h ends using as p 0.21.0. T immed eads we e mapped o he Sewa d ch omosome assembly (monSE13F37) using Page 5 o 21Tahami e al. Mobile DNA (2024) 15:18 BWA mem 0.7.17 [76] wi h de aul pa ame e se ings. PCR duplica es we e emo ed om so ed bam iles wi h sambamba 0.7.0 [77]. Bam iles we e pa sed in o he eebayes 1.3.6 [78] as inpu o a ian calling wi h no popula ion p io , using wo bes SNP alleles, minimum mapping quali y o 20, minimum co e age o 5 and he a 0.02. The VCF iles we e no malized based on he e e - ence genome in 0.57721 [79]. Eigen ec o and eigen- alues we e c ea ed using Plink 2.00a3 [80] a e SNPs we e p uned o linkage disequilib ium (--indep-pai wise 50 10 0.1). We ini ially iden i ied 7,252,950 single nucleo- ide a ian s o he whole genome and 1,773,239 single nucleo ide a ian s wi hin he TE egions. A e he pos - p ocessing, 6,133,447 high quali y SNPs we e eco e ed om he whole genome, om which 6,130,880 we e a iable wi h 6,128,313 (99.9%) biallelic and 2,567 (0.1%) mul iallelic si es. And om he TE egion, 1,342,275 high quali y SNPs we e eco e ed om which 1,340,930 we e a iable wi h 1,339,585 (99.9%) biallelic and 1,345 (0.1%) mul iallelic si es. TE con en The landscapes o TEs using Kimu a 2-pa ame e dis- ance we e plo ed using pe l sc ip s in eg a ed in Repea Maske 4.1.0 using he ch omosome assemblies. To explo e TE abundance di e ences be ween genomes, Chi-squa e es (χ2), using he chisq. es () unc ion in R, was pe o med on he numbe o inse ions in each o he genomes and he Pea son’s esiduals we e plo ed. To al TE densi y was calcula ed be ween he e och o- ma in and euch oma in by in e sec ing masked TEs wi h he iden i ied eu/he e och oma in coo dina es (bed ools .2.30.0 [81]). We me ged masked TEs ha o e lapped in he Repea maske ’s ou pu ile i hey belonged o he same o de , o he wise hey we e named ‘o e lap’ using a cus om sc ip . To calcula e TE densi y ac oss di e en genomic egions, we used he anno a ed g ile o ex ac in ons, exons, ups eam, downs eam and in e genic egions only in he euch oma ic a ea. O e laps be ween egions we e emo ed using bed ools sub ac . We calcula ed he pe cen age o o al leng h o each o he ou TE o de s (LTR, LINE, TIR, RC) pe genomic egion and in euch oma in and he e och oma in, he esul s we e plo - ed using he ggplo 2 package [82]. To es i TEs a e an- domly dis ibu ed ega ding he loca ion o nea by genes, we pe o med Chi-squa e es s a is ics in R using he chisq. es () unc ion o he o al numbe o inse ions o e he size o each compa men . To check i he sequencing and assembly me ics a ec ed he TE con en iden i ied in each genome, we analyzed he co ela ion be ween hese me ics and TE con en es ima es using a linea eg ession model. Iden i ica ion o po en ially ac i e TEs To ind TEs ha a e po en ially ac i e we il e ed he anno a ion ile o TEs ha ha e a leas wo ull leng h copies in he genome. We kep hose TEs ha had equal o o abo e 99% o simila i y ma ch spanning a leas 50% o he o al TE consensus leng h using a cus om bash sc ip . The inal candida es we e manually checked o p o ein coding domains, e minal epea s, and a - ge si e duplica ion (TSD) when applicable, ollowing he me hod desc ibed in Va gas-Cha ez e al. [83]. To calcula e he popula ion equency o hose ac i e TE inse ions, we used PoPoola ionTE2 1.10.03 [84] and TEMP2 0.1.4 pipelines [85] since di e en pipelines migh iden i y di e en TE inse ions. The TE inse ions ha we e de ec ed by bo h PoPoola ionTE2 and TEMP2 pipelines we e me ged using bed ools me ge (bed ools .2.30.0) allowing 25bp dis ance (-d 25) o each inse - ion posi ion. Cold ole an genes We selec ed a o al o 26 cold ole an associa ed genes using p e ious s udies ela ed o D. mon ana cold ole - ance and he published D. mon ana genome assembly based on sho eads [37–43]. We included he genes in o ou cold ole an candida e gene lis i hey we e disco - e ed in a leas h ee di e en s udies (Supplemen a y Table S7). To ans e he coo dina es o ou genome, we pe o med a blas sea ch o he candida e genes in he Sewa d ch omosome-le el assembly. Finally, he new coo dina es o he genes we e checked o homolo- gies in D. melanogas e genome using Flybase e sion FB2023_01 [86]. Each gene was hen scanned up o 1.5kb up and down- s eam o ind TEs. We ocused on TEs ha we e ound o be p esen in all i e genomes analyzed. To check whe he he selec ed TEs could be adding egula o y sequences, we sea ched o ansc ip ion ac o binding si es (TFBS) and p omo e mo i s in TEs up- o down- s eam o he cold ole an genes. We downloaded TFBS mo i s ela ed o s ess esponse in D osophila including immune esponse, hea shock [87, 88], diapause [89–91] and cold shock domain ac o s [92], om JASPAR da a- base [93] (Supplemen a y Table S8). We used da a om D. melanogas e when a ailable and o he wise om Homo sapiens. TFBS we e iden i ied using he web e - sion o FIMO [94] om he MEME SUITE [95, 96] wi h de aul pa ame e s. The ElemeNT online ool was used o iden i y p omo e mo i s [97, 98]. We used R package gggenomes o isual ep esen a ion o ma ched TFBS and p omo e s on he a o emen ioned TEs [99]. Popula ion analysis A hund ed and one wild-caugh , indi idually sequenced emales we e used o es ima e popula ion equencies o Page 6 o 21Tahami e al. Mobile DNA (2024) 15:18 ac i e TEs and TEs loca ed nea by cold ole an genes (Supplemen a y Table S3). A o al o 13 popula ions wi h a minimum numbe o 3 indi iduals pe popula ion we e analyzed using PoPoola ionTE2 1.10.03 [84] and TEMP2 0.1.4 [85]. Raw pai ed-end eads we e immed using as p 0.21.0 wi h minimum quali y Ph ed sco e ≥ 20 (-q 20) and de aul pa ame e s. PoPoola ionTE2 PoPoola ionTE2 allows he de ec ion o e e ence and de no o TE inse ions in genomes. Following he manual, we i s c ea ed a “TE-me ged- e e ence” o D. mon- ana, which consis s o he espec i e masked e e ence genome (monSE13F37) and he TE consensus sequences gene a ed in his wo k. Nex , we c ea ed he “TE hie - a chy ile” by using an ad hoc bash sc ip . T immed aw eads o each sample we e mapped o he co espond- ing TE-me ged- e e ence by using he local alignmen algo i hm BWA bwasw . 0.7.18 ( 1243) [100]. Bo h end eads we e mapped sepa a ely o he TE-me ged- e e - ence, and he pai ed end in o ma ion was es o ed subse- quen ly wi h module se2pe o PoPoola ionTE2. A ppileup was gene a ed o each sample wi h he PoPoola ionTE2 ppileup unc ion (--map-qual 15). Finally, TE inse ions we e iden i ied wi h unc ions iden i ySigna u es (--min- coun 2), he inal se o inse ion pe sample was iden- i ied wi h unc ion pai upSigna u es and ou pu s we e in e sec ed wi h TE coo dina es loca ed nea by cold ol- e an genes using bed ools in e sec . TEMP2 To de ec de no o TE inse ions, we used he module inse ion o TEMP2. We ook he Repea Maske anno a- ion ha was un o PoPoola ionTE2 and ans o med i o bed o ma using g 2bed implemen ed in BEDOPS .2.4.39 [101]. Following he TEMP2 manual, we used BWA mem wi h op ions -Y and -T20 o map he pai ed- end eads o he co esponding e e ence genome. To ge he agmen leng h o he pai ed-end eads, we used Pica d’s Collec Inse SizeMe ics module .2.26.11 [102] and inse ed he mean inse size o each sample. Then, we used TEMP2 inse ion module wi h pa ame e -m 5 (pe cen age o misma ch allowed when mapping o TEs) o de ec TE inse ions. To de ec e e ence inse ions, we used he absence module which anno a es he absence o e e ence TE copies in he samples (pa ame e -x 30, he minimum sco e di e ence be ween he bes hi and he second bes hi o conside ing a ead as uniquely mapped). We con- side ed inse ions ha a e in he e e ence genome and no anno a ed by he absence module as p esen in ou samples. We iden i ied hese inse ions using bed ools in e sec wi h - op ion. Finally, we joined he e e - ence inse ions wi h he ones de ec ed by he inse ion module. Ou pu s we e in e sec ed wi h TE coo di- na es loca ed nea by cold ole an genes using bed ools in e sec . Combining PoPoola ionTE2 and TEMP2 TE inse ions To ind a eliable se o TEs, we combined he TE inse - ions de ec ed by bo h so wa es. We used bed ools wi h op ions me ge -i and -d 25 o collapse inse ions o e lap- ping o allowing a maximum dis ance o 25bp be ween he wo coo dina es in o a single call. We combined inse ions ha belonged o he same TE amily. Ch omosomal in e sions B eakpoin s o one ixed and wo polymo phic D. mon- ana in e sions we e ob ained om Poikela e al. [46] and Poikela e al. (in p ep.) (Supplemen a y Table S9). B ie ly, bo h long- and sho - eads we e used o iden i y he in e sion b eakpoin s. PacBio and Illumina eads we e mapped agains Sewa d genome assembly (monSE13F37) using ngml 0.2.7 [103] and BWA mem, espec i ely, and he esul ing bam iles we e pa sed in o Sni les 2.0.7 [103] and delly 1.1.6 [104] s uc u al a ian iden i i- ca ion p og ams, espec i ely. SURVIVOR 1.0.6 [105] was used o iden i y in e sions ha we e sha ed by bo h s uc u al a ian iden i ica ion p og ams. Finally, he pu a i e b eakpoin s o all in e sions we e con i med isually wi h he IGV using bo h long- and sho - ead da a. To check whe he TEs we e p esen in he in e sion b eakpoin , 50kb egions lanking each side o he b eak- poin s we e mapped agains all i e sca old assemblies using minimap2 o iden i y he b eakpoin coo dina es in each assembly. To alida e he p esence/absence o each b eakpoin , we iden i ied long eads spanning he b eak- poin s (3kb on each side o he b eakpoin o he wo la ge in e sions and 500bp on each side o he sho es in e sion) using IGV. The p oximal and dis al posi ions a e based on hei dis ance o he cen ome e [46]. Resul s Th ee new genome assemblies om D. mon ana eco- geog aphically dis inc popula ions Th ee new D. mon ana e e ence genomes om h ee clima ically and geog aphically di e gen popula ions, C es ed Bu e in No h Ame ica (NA), Oulanka in No h- e n Eu ope (NE) and Kamcha ka in Fa Eas Asia (FE), we e gene a ed in his s udy (Fig.1; Table1, and Supple- men a y Table S2). We also used he wo o he a ailable genome assemblies, Sewa d and Jackson om No h Ame ica D. mon ana popula ions [46]. As sequencing echnology g ea ly impac s TE de ec ion in he genome, we ook ad an age o long- ead sequencing echnique o gene a e hese new assemblies and hus e ie e highe numbe s o TEs compa ed o hose iden i ied based only Page 7 o 21Tahami e al. Mobile DNA (2024) 15:18 on sho - eads [106]. All i e genomes we e sequenced using bo h PacBio and Illumina echnology and assem- bled using he same me hod. In No h Ame ica, he popula ions o Jackson (monJX13F48) and C es ed Bu e (mon34CC5) a e adap ed o ela i ely high la i udes and high ele a ions (1857m and 2960m espec i ely) in he moun ainous egion, whe eas Sewa d (monSE13F37) is a high-la i ude coas al popula ion (35m). In No h Eu ope, Oulanka popula ion (monOU13F149) is loca ed in no h- e n Finland, abo e he a c ic ci cle, and is adap ed o cold and da k win e s and sho summe s. The Asian sample (monKR1309) is om Kamcha ka Peninsula which is a moun ainous olcanic egion wi h long cold win e s in Russia (Fig.1). Upon de no o assembly, he genome sizes o D. mon- ana anged om 175 o 184 Mb. By using a hyb id assembly s a egy, we iden i ied mo e han 95% comple e single copy BUSCOs, wi h N50 alues a ying be ween 0.2 and 11.0Mb (Table1). A e sca olding, he o al genome size and BUSCO alues dec eased (90.6–93.1%, 145–151 Mb) compa ed o he non-sca olded assem- blies. This is because we we e unable o assign all con igs o ch omosome sca olds and he unassigned con igs we e excluded om he inal ch omosome-le el D. mon ana genome [46]. The de ailed in o ma ion on aw eads, con igs and sca olds s a is ics a e gi en in Table1. We used comple e genome assembly o TE iden i ica ion and ch omosome le el assemblies excluding he he e o- ch oma ic egions o he es o ou analyses in his wo k (see Me hods). O e 90% o he TE consensus sequences iden i ied in D. mon ana belong o new amilies o new sub amilies By using genomes om clima ically and geog aphically di e ged D. mon ana popula ions, we wan ed o include as much in aspeci ic di e si y in o ou comp ehensi e TE lib a y as possible. We used REPET o de no o anno- a e TEs in each o he i e genomes. On a e age, 3,221 TE consensus we e buil pe genome, and 303 consensus pe genome we e kep a e manual cu a ion (Supple- men a y Table S10). A e clus e ing indi idual lib a ies, a o al o 555 non- edundan TE consensus sequences we e eco e ed (TE lib a y is p o ided as Addi ional ile 4). 92% o he TE consensuses a e desc ibed as new Table 1 Raw ead and assembly s a is ics o i e D. mon ana genomes gene a ed by PacBio sequencing. N50 aw eads = he ead leng h a which hal o he bases a e in eads longe han o equal o his alue. Raw eads me ics o Sewa d and Jackson a e aken om Poikela e al. [46] Popula ion C es ed Bu e Kamcha ka Oulanka Sewa d Jackson assembly name mon34CC5 monKR1309 monOU13F149 monSE13F37 monJX13F48 Raw eads Raw da a (Gb) 7.2 5.5 8.3 9.9 13.9 Max ead leng h (bp) 145,085 132,961 118,876 168,303 111,203 A e age ead leng h (bp) 5607.9 6268.7 5830 5905.9 6422.3 Mean co e age 39.9 30.6 46.4 54.8 77.1 N50 aw eads 10,260 9233 8836 11,366 8045 De no oassemblies Assembly leng h (Mb) 175.0 176.7 178.9 184.3 181.0 Numbe o con igs 739 1675 1097 324 796 Longes con ig (Mb) 4.5 1.2 3.7 29.1 15.2 N50 (Mb) 0.8 0.2 0.5 11.0 1.3 N50 coun 55 253 97 5 20 GC le el (%) 0.402 0.401 0.401 0.402 0.402 To al comple e BUSCOs (%; n: 3285) 97.2 96 96.4 98.1 98.5 Single copy BUSCOs (%) 96.7 95.6 96.0 97.6 98.0 Duplica ed BUSCOs (%) 0.5 0.4 0.4 0.5 0.5 Sca old assemblies Assembly leng h (Mb) 148.2 148.7 149.8 145.5 151.8 Numbe o sca olds 6 6 6 6 6 Sca old leng h (Mb) 34.8 34.4 33.3 32.5 33.9 N50 (Mb) 26.6 26.5 26.7 26.5 28.4 N50 coun 3 3 3 3 3 GC le el (%) 0.402 0.402 0.401 0.403 0.402 To al comple e BUSCOs (%; n: 3285) 92.4 90.6 90.6 91.9 93.1 Single copy BUSCOs (%) 92.2 90.4 90.4 91.7 92.8 Duplica ed BUSCOs (%) 0.2 0.2 0.2 0.20 0.3 Accession numbe SAMN27782986 SAMN27782984 SAMN27782985 SAMN27782981 SAMN27782982 Page 8 o 21Tahami e al. Mobile DNA (2024) 15:18 amilies (445) and new sub amilies (66) and only he emaining 8% co esponds o p e iously known ami- lies (44) (Fig. 2 and Supplemen a y Table S11). Long e minal epea s (LTR) a e he mos abundan TE o de in he D. mon ana lib a y o which, Gypsy supe amily is he mos abundan (80%) ollowed by Bel-Pao (14%). LINE and DNA ansposons (TIR, e minal in e ed epea , and RC, olling ci cle elemen s) con ibu e almos equally o he TE lib a y (12%) wi h he highes abun- dance o Jockey (68%) and Tc1-Ma ine (40%) in each o de , espec i ely (Fig.2). We used he manual cu a ed lib a y o anno a e each o he i e e e ence genomes and analyzed he abundance, dis ibu ion, and ac i i y o TE sequences. We i s analyzed whe he di e ences in sequencing o assembly me ics explained a ia ion in he TE me ics es ima ed (Supplemen a y Table S12). The N50 o con igs had a signi ican e ec on he o al numbe o aw TE sequences disco e ed in he he e o- ch oma in, while sca old s a is ics did no impac any o he TE me ics analyzed (Supplemen a y Table S12). Because all o he ollowing analyses ha e been made on he euch oma in, ou esul s a e hus no a ec ed by sequencing o assembly di e ences ac oss genomes. We also explo ed he gene ic a ia ion ac oss he whole genome and wi hin TE sequences in he i e D. mon ana samples by pe o ming a p incipal componen analysis (PCA) wi h biallelic SNPs. Resul s om he wo PCA analyses a e highly conco dan , wi h he i s PC sepa- a ing he h ee NA popula ions om he Oulanka (NE) and Kamcha ka (FE) popula ions and explaining simila amoun s o a ia ion: 33 and 32% o he whole genome and TE egion SNPs, espec i ely (Fig.3A and B and Sup- plemen a y Table S13). The second PC, which explains 28 and 29% o he a ia ion, sepa a es he C es ed Bu e popula ion (NA) om he o he wo NA popula ions: Jackson and Sewa d. This pa e n likely e lec s he demog aphy o hese popula ions (Poikela e al., in p ep) and shows he high le el o in aspeci ic a ia ion in he D. mon ana genomes analyzed. T ansposable elemen s con ibu e 11–13% o he D. mon ana genomes Whole genome TE con en ac oss genomes a ied be ween 10.7% and 13% (Table2). In all i e genomes, LTRs we e he mos abundan o de , ollowed by olling ci cles (RC), LINEs, and e minal in e ed epea s (TIRs) (Fig. 4A). TE abundance a he o de and supe amily le el a ied ac oss genomes (Fig.4B and C). In Sewa d, LTR (χ2 esidue − 8.2, p- alue < 0.05) is he leas abundan and RC (χ2 esidue 5.7, p- alue < 0.05) is he mos abun- dan o de compa ed wi h he o he genomes analyzed. A he supe amily le el, Sewa d shows he highes a ia- ion in TE con en ; Gypsy and Bel-Pao ha e less inse - ions compa ed wi h o he genomes (χ2 esidue − 11.1 Fig. 2 Consensus sequence classi ica ions o he cu a ed TE lib a y buil om D. mon ana genomes. (A) Pe cen age o new o known TE consensus (cons) sequences. (B) Numbe s o consensus elemen s classi ied as LTR, DNA (RC and TIR) and LINEs Page 15 o 21Tahami e al. Mobile DNA (2024) 15:18 Table 3 T ansposable elemen s ha we e ound in he b eakpoin egions (3kb/500bp) o he h ee in e sions analyzed in his wo k. TEs ha a e p esen only in he genomes wi h in e sion b eakpoin s (and no in o he genomes) a e in bold Ch In e sion coo dina es Genomes TE con en P oximal Dis al P oximal b eakpoin Dis al b eakpoin X 11,174,369–11,174,380 3,994,112–3,994,127 Sewa d R1-2_Dmon Ma e ick-1_Dmon TART_DV-Dmon-B Heli on-5_Dmon Gypsy-102_Dmon Heli on-10_Dmon Heli on-2N1_D i -Dmon-B Heli on-2N1_D i -Dmon-A Heli on-1N1_D i 11,136,142–11,136,155 3,936,207–3,936,246 Jackson R1-2_Dmon Heli on-5_Dmon Gypsy-102_Dmon Heli on-10_Dmon Heli on-5_Dmon Heli on-2N1_D i -Dmon-B Heli on-2N1_D i -Dmon-A Heli on-1N1_D i 11,110,783–11,110,796 3,984,908–3,984,929 Kamcha ka R1-2_Dmon Ma e ick-1_Dmon Heli on-5_Dmon Gypsy-102_Dmon Heli on-10_Dmon Heli on-5_Dmon Heli on-2N1_D i -Dmon-B Heli on-2N1_D i -Dmon-A Heli on-1N1_D i 11,304,271–11,304,284 4,034,972–4,035,011 Oulanka Ma e ick-1_Dmon Heli on-5_Dmon Gypsy-102_Dmon Heli on-10_Dmon Heli on-8_Dmon Heli on-2N1_D i -Dmon-B Heli on-2N1_D i -Dmon-A Heli on-1N1_D i Ma e ick-1_Dmon Ma e ick-1_Dmon 10,289,870–10,289,883 3,549,579–3,549,600 C es ed Bu e R1-2_Dmon Heli on-5_Dmon Gypsy-102_Dmon Heli on-10_Dmon Heli on-5_Dmon Heli on-1N1_D i Heli on-2N1_D i -Dmon-B Heli on-2N1_D i -Dmon-A Heli on-1N1_D i Ma e ick-1_Dmon 4 6,596,948–6,599,447 7,298,775 − 7,298,610 C es ed Bu e Gypsy-5_DVi -I-Dmon-F Gypsy-40_Dmon Gypsy-5_DVi -I-Dmon-A MITE-1_Dmon MITE-1_Dmon Heli on-1_D i -Dmon-C Gypsy-215_Dmon Gypsy-181_Dmon 4 2,906,226 2,907,786–2,907,896 Oulanka Heli on-4_Dmon Gypsy-83_Dmon Gypsy-70_Dmon Heli on-5_Dmon Gypsy-8_DVi -I-Dmon-B ULYSSES_I-Dmon-H Heli on-4_Dmon Heli on-1_D i -Dmon-C 3,165,124–3,165,142 3,166,689 Kamcha ka Heli on-4_Dmon Gypsy-83_Dmon Gypsy-70_Dmon Heli on-5_Dmon Gypsy-8_DVi -I-Dmon-B Heli on-4_Dmon Page 16 o 21Tahami e al. Mobile DNA (2024) 15:18 o de and Gypsy supe amily (Fig. 2). Only 8% o he TE consensus sequences co esponded o known TE amilies, mos ly o D. i ilis anno a ions, which is mo e closely ela ed o he s udied species han D. melanogas- e (Supplemen a y Table S11). The e a e se e al po en- ial easons o he high numbe o new TE amilies. P obably he mos impo an eason is ha D. mon ana is highly di e ged om he o he model and non-model D osophila species o which TE lib a ies exis in da a- bases (be ween ~ 40 MYA om D. melanogas e and 9 MYA om D. i ilis) [112, 113]. Ano he plausible eason comes om he species’ unique dis ibu ion in he Hol- a c ic egion. Upon coloniza ion in o new habi a s, spe- cies can be exposed o he in asion o new TEs h ough ho izon al ans e om new pa asi es and om o he conspecies [114]. Simila e en s could ha e happened in D. mon ana while occupying new cold habi a s. Ou p incipal componen analysis (PCA) on genome- wide dis ibu ion o SNPs and wi hin TE egions shows he high le el o in aspeci ic a ia ion ha we aimed o include in o ou TE lib a y (Fig.3). The PCA depic s he high di e gence no only be ween D. mon ana popula ions om No h Ame ica and No he n Eu ope, bu also among No h Ame ican popula ions as C es ed Bu e (Colo ado) is highly di e ged om Jackson (Wyo- ming) and Sewa d (Alaska). High di e gence o C es ed Bu e and o he No h Ame ican popula ions is likely a consequence o pas ounde e ec s po en ially associ- a ed wi h gene ic d i , a iable selec ion p essu es, and local adap a ion [43]. Mo eo e , he la ge in e sion on ch omosome 4 is uniquely ixed in C es ed Bu e (Poikela e al., in p ep.), which u he educes gene ic exchange be ween popula ions and esul s in ele a ed gene ic di e gence (Poikela e al., in p ep. and [27]). O e all, en i onmen al he e ogenei y h ough ime and space can ac as a selec i e o ce d i ing adap i e di e en ia ion be ween popula ions. In e es ingly, Kamcha ka shows high simila i y o he Finnish popula ion om Oulanka ega dless o hei geog aphic dis ance. This a ini y could be explained by he ances al ou e o expansion om NA popula ions owa ds Eu asia h ough he Be ing S ai [35], he e o e, he gene ic di e gence o Eu opean and Asian popula ions would be low. Fig. 8 Polymo phic and ixed in e sion b eakpoin s in D. mon ana. Fo each in e sion bo h b eakpoin s, p oximal (close o he cen ome e) and dis al ( a he om he cen ome e), plus ~ 3kb o 500bp on each side a e shown. When he posi ion o a b eakpoin was no iden i ied a he single base pai le el, he in e al whe e he b eakpoin is p edic ed o be is shown in a g ey box. Genes a e shown as da k blue boxes, and TEs a e colo ed based on hei supe amily le el. B eakpoin posi ions a e d awn based on hei same posi ion on he aligned ch omosomes in he sense di ec ion Page 17 o 21Tahami e al. Mobile DNA (2024) 15:18 TE con en in D. mon ana Ou de no o TE anno a ions con i med ha he amoun o TE con en signi ican ly con ibu es o he D. mon ana genome size as epo ed o o he species [10, 115]. The o e all a e age TE con en in D. mon ana, 11–13%, is simila o ha o i s close ela i es, D. i ilis (15%) [11] and D. incomp a (13–14%) [107], and smalle han hose o D. melanogas e (20%) [115] and D. suzukii (30%) [11]. LTRs ep esen he highes genomic ac ion o ~ 5% ol- lowed by RC elemen s (Heli on and Ma e ick), ~ 4%, LINE, ~ 2%, and inally TIR wi h < 1% (Fig.4A). Acco d- ing o his obse ed pa e n, LTRs a e he mos abundan elemen s which is ypical o many D osophila lies, how- e e , he abundance does no ollow he assumed global pa e n: LTRs > LINEs > TIRs > OTHERs [10, 11]. The highe Heli on ac ion in D. mon ana genome (com- pa ed o he global pa e n) combined wi h i s a he pe sis en ac i i y spanned o e an ex ended cou se o ime (Fig.6A) sugges s ha Heli ons ha e been ole - a ed mo e han o he elemen s p obably due o hei neu al o adap i e e ec s. The elaxed selec i e p essu e on Heli on ansposi ion could come om he elemen ’s own na u e and i s endency o inse nea less c i ical genes. Howe e , esul s om phylogene ically ela i ely close ela i es o D. mon ana such as D. buzza ii and D. moja ensis, ha e indica ed he abundance o Heli on in hei genomes, con ibu ing o he hypo hesis ha Heli- ons can ha e a wide dis ibu ion in he subgenus D o- sophila [116]. In D. mon ana a signi ican ly highe numbe o TEs we e accumula ed in he he e och oma in egion. The he e och oma in’s esilience o TE accumula ion is due o he sca ci y o ac i e genes, and as expec ed, TEs we e less ole a ed in he euch oma in (Fig.5A). Acco ding o he D osophila 12 Genomes Conso ium (2007), 1–9% o he euch oma in is occupied by TEs ac oss D osophila species, and hence, he calcula ed 7% o D. mon ana is wi hin he epo ed ange. Analysis o he TE con en dis ibu ion be ween di e en genomic egions e eals highe accumula ion o TEs in he in e genic a ea ol- lowed by in ons, wi h he leas occupancy in exons (Fig. 5D). O e all, TE dis ibu ion was no andom h oughou he genome o be ween ch omosomes sug- ges ing ha pu i ying selec ion ac s agains TE inse ions due o hei mul iple dele e ious e ec s [117]. TE den- si y was signi ican ly highe in ch omosome 4 and X and he lowes in he igh a m o ch omosome 2 (2R). Apa om he ch omosome sizes, he highe acquisi ion o TEs on sex ch omosome could be associa ed wi h dosage compensa ion [118]. Mo eo e , Wibe g e al. [43] ound ha he op candida e SNPs o cold adap a ion in D. mon ana a e en iched on ch omosomes 4 and X which also include se e al in e sions (see below). Acco ding o he landscape o TEs, all i e genomes show simila dynamics o he TE con en . We ound ha he D. mon ana TE landscape is compa able o ha o D. melanogas e which also has a la ge ac ion o young and ac i e LTR elemen s, while i di e s om he D. simulans landscape which ins ead has a la ge ac ion o old and deg aded elemen s [11]. The pa e n is bimodal wi h a gene al end owa ds a small numbe o highly di e ged/ old elemen s (Fig.6A). The dynamic pa e n o TE land- scapes is o en a e lec ion o he species demog aphic his o y [119, 120]. Fo example, occupying new en i on- men s will pose a ious ypes o s ess o he species o which, cold is one s ess ac o in he no he n egions, which can po en ially e-ac i a e mobile elemen s [12, 121]. The bimodal pa e n indica es wo bu s s o ans- posi ion e en s ha ha e in e up ed he ansposi ion/ excision equilib ium (s anda d bell shape pa e n). This pa e n is gene ally seen in specialis species and spe- cies wi h small e ec i e popula ion size which happens when he dis ibu ion is pa chy and limi ed by en i on- men al esou ces [107]. The i e genomes also ha e simi- la numbe s o po en ially ac i e TEs (Fig.6B), which a e mos ly p esen a low popula ions equencies (Fig.6C) as expec ed i hese a e young inse ions. TEs associa ed wi h cold ole an genes Ou analysis o he TEs loca ed nea by cold- ole an genes in he i e genomes sugges ha some o hem migh be playing a ole in he egula ion o hese genes as hey a e p esen a high popula ion equencies and con ain egula o y sequences (Fig. 7 and Supplemen- a y Table S21). Th ee ypes o FOXO1 and one BEAF- 32 TFBS, among o he TFBS, we e ound inside some o hese TEs (Fig.7). Heli on-2N1_D i -Dmon-B loca ed downs eam o Yp2 and ups eam o CG12057 also has p omo e mo i s and FOXO ansc ip ion ac o s. FOXO ansc ip ion ac o s egula e cellula homeos asis, lon- ge i y and esponse o s ess [122]. In D. mon ana, ac i- a ion o FOXO has been sugges ed o be connec ed wi h he lies’ o e win e ing abili y [89]. Mo eo e , Heli- on-1_D i -Dmon-C ups eam o CG17571 has a ma ch o p omo e elemen s only in he Kamcha ka genome. This TE ma ches BEAF-32 TFBS which is a ch oma in insula o and a ec s he egula ion o gene ansc ip- ion [88]. BEAF-32 is also connec ed o he de elopmen o pho o ecep o di e en ia ion du ing emb yogenesis [91]. The e o e, hese candida e TEs ha e he po en ial o a ec he egula ion o he lanking cold ole an genes, as has been shown be o e o D. melanogas e TE inse - ions [87]. TEs associa ed wi h in e sions TEs and o he epe i i e egions ha e been sugges ed o be in ol ed in he o igin o ch omosomal in e sions Page 18 o 21Tahami e al. Mobile DNA (2024) 15:18 [123]. By s udying h ee in e sions ecen ly cha ac e ized in D. mon ana ac oss i s dis ibu ion ( [46] and Poikela e al., in p ep), we we e able o sea ch o he p esence o TEs nea in e sion b eakpoin s. These in e sions we e ound on ch omosomes X and 4, which also ha e he highes TE densi y. Pe haps no su p isingly, we ound TEs nea by all he b eakpoin s. The b eakpoin s o he la ge ch omosome 4 in e sion ha is ixed in C es ed Bu e ha bo s i e TE inse ions ha a e exclusi e o C es ed Bu e and no ound in he o he popula ions. Simila ly, he b eakpoin s o he small ch omosome 4 in e sion ound in Oulanka and Kamcha ka show TE inse ions ha a e speci ic o hese popula ions bu a e no ound in he o he s. These indings could sugges ha some o hese TE inse ions ound a ound he in e - sion b eakpoin s may ha e ac ed as subs a es o ec opic ecombina ion and acili a ed he o igin o he in e sions, al hough u he analysis would be needed o con i m i [123]. Finally, he b eakpoin s o he X in e sion ha e se e al TE inse ions ha a e p esen in all D. mon ana popula ions, in acco dance wi h he in e sion being ixed in D. mon ana. Some o hese inse ions could ha e po en ially played a ole in he es ablishmen o ha X in e sion ~ 3 MY [46], bu would equi e an in-dep h in es iga ion in o he TEs be ween D. mon ana and o he species o he mon ana phylad (D. la omon ana, D. bo ealis, D. lacicola). O e all, hese TE inse ions can g ea ly impac he o ma ion o la ge ea angemen s, po en ially leading o signi ican e olu iona y ou comes. Fo example, he la ge ch omosome 4 in e sion is asso- cia ed wi h SNPs linked o cold/clima e adap a ion in D. mon ana and wi h a educ ion in gene exchange be ween D. mon ana popula ions (Poikela e al., in p ep.). Addi- ionally, he X in e sion is linked o he specia ion p o- cess be ween D. mon ana and D. la omon ana [46, 124]. Conclusions S udying TEs in non-model species expands ou unde - s anding o TE dynamics h oughou zoo axa, and hei a ious impac s on he genome e olu ion. Adding ca e- ully cu a ed and a non- edundan TE lib a y om non- model species o he public eposi o ies signi ican ly con ibu es o his mission. We eco e ed 555 cu a ed non- edundan TE consensuses wi h ema kable num- be s o new amilies and sub amilies comp ising o e 90% o he TE lib a y in D. mon ana, using gene ically dis- an in aspecies samples. TE con ibu ion o he whole genome con en is gene ally smalle in D. mon ana han i s ela i e D osophila species assuming ha di e ences in he quali y o he genome assemblies be ween di e - en species a e no he case. As expec ed TEs we e non- andomly dis ibu ed, we also ound a highe densi y in he e och oma ic egions, and he X ch omosome com- pa ed o au osomes. Howe e , he TE abundance a he o de le el seems o a y ac oss di e en species which equi es mo e in es iga ion om he e olu iona y s and- poin . Mo eo e , we ound TEs nea by o in he i s in on o a selec ed se o p e iously iden i ied cold ol- e an genes which ha e ansc ip ion ac o binding si es and/o p omo e elemen s, indica ing hei po en ial o in luencing gene’s egula ion. Howe e , ou cu en a emp was only di ec ed o a se o genes iden i ied in ou p e ious s udies. Undoub edly, u u e esea ch will disco e mo e cold ela ed genes and TEs connec ed wi h hem. Finally, o he i s ime, we ound se e al TEs a ound he h ee in e sion b eakpoin s in D. mon ana, which a e wo h u he s udies on hei po en ial impac on in e sions. Abb e ia ions TE T ansposable elemen NA No h Ame ica NE No h Eu ope FE Fa Eas TFBS T ansc ip ion ac o binding si es LTR Long e minal epea s LINE Long in e spe sed nuclea elemen s TIR Te minal in e ed epea s RC Rolling ci cles Supplemen a y In o ma ion The online e sion con ains supplemen a y ma e ial a ailable a h ps://doi. o g/10.1186/s13100-024-00328-7. Supplemen a y Ma e ial 1 Supplemen a y Ma e ial 2 Supplemen a y Ma e ial 3 Supplemen a y Ma e ial 4 Acknowledgemen s The au ho s would like o hank MSc Sa a Lommi o he wo k in he labo a o y and in he ly s ain main enance. Analyses we e ca ied ou using CSC (Finnish IT Cen e o Science) se ices. Au ho con ibu ions MST: W o e he manusc ip , p epa ed he ables and igu es, de eloped code and ca ied ou he analyses, CVC: Con ibu ed o analyses and p oo eading. NP: Con ibu ed o analyses, w i ing, and p oo eading. MCZ: Con ibu ed o analyses, and igu es, and de eloped code, p oo eading. JG: Con ibu ed o concep ualiza ion, w i ing, and p oo eading. MK: concei ed and designed he p ojec , con ibu ed o w i ing, p oo eading, and p ojec managemen . All he au ho s ead and app o ed he inal manusc ip . Funding JG is unded by g an PID2020-115874GB-I00 awa ded by MICIU/AEI/h ps:// doi.o g/10.13039/501100011033/ and om g an 2021 SGR 00417 awa ded by Depa amen de Rece ca i Uni e si a s, Gene ali a de Ca alunya awa ded o J.G. MK was unded by he Academy o Finland p ojec 322980. The unde s had no ole in he p epa a ion and publica ion o his manusc ip . The es o he au ho s ha e decla ed no unding associa ed wi h his wo k. Da a a ailabili y The da ase s suppo ing he conclusions o his a icle a e included wi hin he a icle and i s addi ional iles. The genome assemblies a e e ie able wi h p ojec name PRJNA828433 and p e iously published assemblies and PacBio and Illumina aw eads a e a ailable unde Biop ojec PRJNA939085. The TE Page 19 o 21Tahami e al. Mobile DNA (2024) 15:18 lib a y is p o ided as Addi ional ile 4. The sc ip s de eloped o he pu pose o his s udy a e a ailable in he Gi hub eposi o y, h ps://gi hub.com/ ahami-ms/TE-p ojec . Decla a ions E hics app o al and consen o pa icipa e No applicable. Consen o publica ion No applicable. Compe ing in e es s The au ho s decla e no compe ing in e es s. Recei ed: 9 Ap il 2024 / Accep ed: 17 Sep embe 2024 Re e ences 1. Cohen J, Agel L, Ba low M, Ga inkel CI, Whi e I. Linking A c ic a iabili y and change wi h ex eme win e wea he in he Uni ed S a es. Sci (1979). 2021;373:1116–21. 2. Coumou D, Rahms o S. A decade o wea he ex emes. Na Clim Chang. 2012;2:491–6. 3. Sp ingael D, Top EM. Ho izon al gene ans e and mic obial adap a ion o xenobio ics: new ypes o mobile gene ic elemen s and lessons om eco- logical s udies. T ends Mic obiol. 2004;12:53–8. 4. Ki kpa ick M, Ba on N. Ch omosome in e sions, local adap a ion and specia- ion. Gene ics. 2006;173:419–34. 5. Sch ade L, Kim JW, Ence D, Zimin A, Klein A, Wysche zki K, e al. T ansposable elemen islands acili a e adap a ion o no el en i onmen s in an in asi e species. Na Commun. 2014;5:1–10. 6. Dudaniec RY, Yong CJ, Lancas e LT, S ensson EI, Hansson B. Signa u es o local adap a ion along en i onmen al g adien s in a ange-expanding dam- sel ly (Ischnu a Elegans). Mol Ecol. 2018. 7. Cayuela H, Do an Y, Mé o C, Lapo e M, No mandeau E, Gagnon-Ha ey S, e al. The mal adap a ion a he han demog aphic his o y d i es gene ic s uc u e in e ed by copy numbe a ian s in a ma ine ish. Mol Ecol. 2021;30:1624–41. 8. de Koning APJ, Gu W, Cas oe TA, Ba ze MA, Pollock DD. Repe i i e ele- men s may comp ise o e wo- hi ds o he human genome. PLoS Gene . 2011;7:e1002384. 9. Vicien CM. T ansc ip ional ac i i y o ansposable elemen s in maize. BMC Genomics. 2010;11:1–10. 10. Pe e sen M, A misén D, Gibbs RA, He ing L, Khila A, Maye G, e al. Di e si y and e olu ion o he ansposable elemen epe oi e in a h opods wi h pa icula e e ence o insec s. BMC Ecol E ol. 2019;19:1–15. 11. Mé el V, Boules eix M, Fable M, Viei a C. T ansposable elemen s in D osophila. Mob DNA. 2020;11:1–20. 12. Ho á h V, Me enciano M, González J. Re isi ing he ela ionship be ween ansposable elemen s and he euka yo ic s ess esponse. T ends Gene . 2017;33:832–41. 13. Maside X, Ba olomé C, Cha leswo h B. S-elemen inse ions a e associa ed wi h he e olu ion o he Hsp70 genes in D osophila melanogas e . Cu Biol. 2002;12:1686–91. 14. Nai o K, Zhang F, Tsukiyama T, Sai o H, Hancock CN, Richa dson AO, e al. Unexpec ed consequences o a sudden and massi e ansposon ampli ica- ion on ice gene exp ession. Na u e. 2009;461:1130–4. 15. Amine zach YT, Macphe son JM, Pe o DA. Pes icide esis ance ia ans- posi ion-media ed adap i e gene unca ion in D osophila. Science (1979). 2005;309:764–7. 16. Salces-O iz J, Va gas-Cha ez C, Guio L, Rech GE, González J. T ansposable elemen s con ibu e o he genomic esponse o insec icides in D osophila melanogas e . Philosophical T ans Royal Soc B: Biol Sci. 2020;375:20190341. 17. Guio L, Viei a C, González J. S ess a ec s he epigene ic ma ks added by na u al ansposable elemen inse ions in D osophila melanogas e . Sci Rep. 2018;8:1–10. 18. Casacube a E, González J. The impac o ansposable elemen s in en i on- men al adap a ion. Mol Ecol. 2013;22:1503–17. 19. Se a o-Capuchina A, Ma u e DR. The ole o ansposable elemen s in specia- ion. Genes (Basel). 2018;9:254. 20. Co onado-Zamo a M, González J. T ansposons con ibu e o he unc ional di e si ica ion o he head, gu , and o a y ansc ip omes ac oss D osophila na u al s ains. Genome Res. 2023;33:1541–53. 21. Me enciano M, Ullas es A, de Ca a MAR, Ba on MG, Gonzalez J. Mul iple independen e oelemen inse ions in he p omo e o a s ess esponse gene ha e a iable molecula and unc ional e ec s in D osophila. PLoS Gene . 2016;12:e1006249. 22. Me enciano M, González J. The in e play be ween de elopmen al s age and en i onmen unde lies he adap i e e ec o a na u al ansposable elemen inse ion. Mol Biol E ol. 2023;40:msad044. 23. Joly-Lopez Z, Bu eau TE. Exap a ion o ansposable elemen coding sequences. Cu Opin Gene De . 2018;49:34–42. 24. Reis M, Viei a CP, La a R, Posnien N, Viei a J. O igin and consequences o ch omosomal in e sions in he i ilis g oup o D osophila. Genome Biol E ol. 2018;10:3152–66. 25. Rius N, Delp a A, Ruiz A. A di e gen P elemen and i s associa ed MITE, BuT5, gene a e ch omosomal in e sions and a e widesp ead wi hin he D osophila eple a species g oup. Genome Biol E ol. 2013;5:1127–41. 26. Delp a A, Neg e B, Puig M, Ruiz A. The ansposon Galileo gene a es na u al ch omosomal in e sions in D osophila by ec opic ecombina ion. PLoS ONE. 2009;4:e7883. 27. Ho mann AA, Riesebe g LH. Re isi ing he impac o in e sions in e olu ion: om popula ion gene ic ma ke s o d i e s o adap i e shi s and specia ion? Annu Re Ecol E ol Sys . 2008;39:21–42. 28. Jackson BC. Recombina ion-supp ession: how many mechanisms o ch o- mosomal specia ion? Gene ica. 2011;139:393–402. 29. Fa ia R, Na a o A. Ch omosomal specia ion e isi ed: ea anging heo y wi h pieces o e idence. T ends Ecol E ol. 2010;25:660–9. 30. Ayala D, Fon aine MC, Cohue A, Fon enille D, Vi alis R, Sima d F. Ch omo- somal in e sions, na u al selec ion and adap a ion in he mala ia ec o Anopheles Funes us. Mol Biol E ol. 2011;28:745–58. 31. E gen’e MB, Zelen so a H, Poluec o a H, Lyozin GT, Veleikod o skaja V, Pya ko KI e al. Mobile elemen s and ch omosomal e olu ion in he i ilis g oup o D osophila. P oceedings o he Na ional Academy o Sciences. 2000;97:11337–42. 32. Th ockmo on LH. The i ilis species g oup. Gene Bioogy D osophila. 1982;3:227–96. 33. MacMillan HA, Ande sen JL, Da ies SA, O e gaa d J. The capaci y o main ain ion and wa e homeos asis unde lies in e speci ic a ia ion in D osophila cold ole ance. Sci Rep. 2015;5:1–11. 34. Poikela N, Tyukmae a V, Hoikkala A, Kanka e M. Mul iple pa hs o cold ole - ance: he ole o en i onmen al cues, mo phological ai s and he ci cadian clock gene ille. BMC Ecol E ol. 2021;21:1–20. 35. Hoikkala A, Poikela N. Adap a ion and ecological specia ion in seasonally a ying en i onmen s a high la i udes: D osophila i ilis g oup. Fly (Aus in). Taylo and F ancis L d.; 2022. pp. 85–104. 36. Mi ol PM, Schä e MA, O sini L, Rou u J, Schlö e e C, Hoikkala A, e al. Phylo- geog aphic pa e ns in D osophila mon ana. Mol Ecol. 2007;16:1085–97. 37. Kanka e M, Salminen T, Laiho A, Vesala L, Hoikkala A. Changes in gene exp es- sion linked wi h adul ep oduc i e diapause in a no he n mal ly species: a candida e gene mic oa ay s udy. BMC Ecol. 2010;10:1–9. 38. Kanka e M, Pa ke DJ, Me isalo M, Salminen TS. T ansc ip ional di e ences be ween Diapausing and Non-diapausing D. Mon ana emales ea ed unde he same pho ope iod and empe a u e. 2016;1–18. 39. Kau anen H, Kinnunen J, Hiillos A-L, Lankinen P, Hopkins D, Wibe g RAW, e al. Selec ion o ep oduc ion unde sho pho ope iods changes diapause- associa ed ai s and induces widesp ead genomic di e gence. J Exp Biol. 2019;222:jeb205831. 40. Pa ke DJ, Ri chie MG, Kanka e M. P epa ing o win e : he ansc ip omic esponse associa ed wi h di e en day leng hs in D osophila mon ana. G3 Genes|Genomes|Gene ics. 2016;6:1373–81. 41. Pa ke DJ, Vesala L, Ri chie MG, Laiho A, Hoikkala A, Kanka e M. How consis- en a e he ansc ip ome changes associa ed wi h cold acclima ion in wo species o he D osophila i ilis g oup? He edi y (Edinb). 2015;115:13–21. 42. Pa ke DJ, Wibe g RAW, T i edi U, Tyukmae a VI, Gha bi K, Bu lin RK, e al. In e and in aspeci ic genomic di e gence in D osophila mon ana shows e idence o cold adap a ion. Genome Biol E ol. 2018;10:2086–101. 43. Wibe g RAW, Tyukmae a V, Hoikkala A, Ri chie MG, Kanka e M. Cold adap a- ion d i es popula ion genomic di e gence in he ecological specialis , D osophila mon ana. Mol Ecol. 2021;30:3783–96. Page 20 o 21Tahami e al. Mobile DNA (2024) 15:18 44. Mo ales-Hojas R, Päällysaho S. Compa a i e poly ene ch omosome maps o D. mon ana. Ch omosoma. 2007;116(1):21–7. 45. S one WS, Gues WC, Wilson FD. The e olu iona y implica ions o he cy o- logical polymo phism and phylogeny o he i ilis g oup o D osophila. P oc Na l Acad Sci. 1960;46:350–61. 46. Poikela N, Lae sch DR, Hoikkala V, Lohse K, Kanka e M. Ch omosomal In e sions and he Demog aphy o Specia ion in D osophila mon ana and D osophila la omon ana. Gonzalez J, edi o . Genome Biol E ol. 2024;3:16. 47. Ruan J, Li H. Fas and accu a e long- ead assembly wi h w dbg2. Na Me h- ods. 2020;17:155–8. 48. Zimin AV, Puiu D, Luo M-C, Zhu T, Ko en S, Ma çais G, e al. Hyb id assembly o he la ge and highly epe i i e genome o Aegilops auschii, a p ogeni o o b ead whea , wi h he MaSuRCA mega- eads algo i hm. Genome Res. 2017;27:787–92. 49. Chak abo y M, Baldwin-B own JG, Long AD, Eme son JJ. Con iguous and accu a e de no o assembly o me azoan genomes wi h modes long ead co e age. Nucleic Acids Res. 2016;44:e147–147. 50. Walke BJ, Abeel T, Shea T, P ies M, Abouelliel A, Sak hikuma S, e al. Pilon: an in eg a ed ool o comp ehensi e mic obial a ian de ec ion and genome assembly imp o emen . PLoS ONE. 2014;9:e112963. 51. Guan D, McCa hy SA, Wood J, Howe K, Wang Y, Du bin R. Iden i ying and emo ing haplo ypic duplica ion in p ima y genome assemblies. Bioin o - ma ics. 2020;36:2896–8. 52. Lae sch DR, Blax e ML, BlobTools. In e oga ion o genome assemblies. F1000Res. 2017;6:1287. 53. Seppey M, Manni M, Zdobno EM. BUSCO: assessing genome assembly and anno a ion comple eness. Gene P edic ion: Me hods P o ocols. 2019;227–45. 54. Alonge M, Lebeigle L, Ki sche M, Jenike K, Ou S, Aganezo S, e al. Au oma ed assembly sca olding using RagTag ele a es a new oma o sys em o high- h oughpu genome edi ing. Genome Biol. 2022;23:1–19. 55. Li H. Minimap2: pai wise alignmen o nucleo ide sequences. Bioin o ma ics. 2018;34:3094–100. 56. Flu e T, Dup a E, Feuille C, Quesne ille H. Conside ing ansposable elemen di e si ica ion in de no o anno a ion app oaches. PLoS ONE. 2011;6:e16526. 57. Bao Z, Eddy SR. Au oma ed de no o iden i ica ion o epea sequence ami- lies in sequenced genomes. Genome Res. 2002;12:1269–76. 58. Malik L, Almoda esi F, Pa o R. G oupe : g aph-based clus e ing and anno a ion o imp o ed de no o ansc ip ome analysis. Bioin o ma ics. 2018;34:3265–72. 59. Edga RC, Mye s EW. PILER: iden i ica ion and classi ica ion o genomic epea s. Bioin o ma ics-Ox o d. 2005;21:i152. 60. Ju ka J, Kapi ono VV, Pa licek A, Klonowski P, Kohany O, Walichiewicz J. Repbase Upda e, a da abase o euka yo ic epe i i e elemen s. Cy ogene Genome Res. 2005;110:462–7. 61. Tho aldsdó i H, Robinson JT, Mesi o JP. In eg a i e Genomics Viewe (IGV): high-pe o mance genomics da a isualiza ion and explo a ion. B ie Bioin o m. 2013;14:178–92. 62. Wicke T, Sabo F, Hua-Van A, Benne zen JL, Capy P, Chalhoub B, e al. A uni ied classi ica ion sys em o euka yo ic ansposable elemen s. Na Re Gene . 2007;8:973–82. 63. O ozco S, Sie a P, Du bin R, Gonzalez J. MCHelpe au oma ically cu a es ansposable elemen lib a ies ac oss species. bioRxi . 2023;2010–23. 64. Smi AFA, Hubley R, G een P. Repea Maske Open-4.0. 2013. 65. Pa ke DJ, En all T, Ri chie MG, Kanka e M. Sex-speci ic esponses o cold in a e y cold- ole an , no he n D osophila species. He edi y (Edinb). 2021;126:695–705. 66. Chen S, Zhou Y, Chen Y, Gu J. Fas p: an ul a- as all-in-one FASTQ p ep oces- so . Bioin o ma ics. 2018;34:i884–90. 67. Dobin A, Da is CA, Schlesinge F, D enkow J, Zaleski C, Jha S, e al. STAR: ul a as uni e sal RNA-seq aligne . Bioin o ma ics. 2013;29:15–21. 68. B ůna T, Ho KJ, Lomsadze A, S anke M, Bo odo sky M. BRAKER2: au oma ic euka yo ic genome anno a ion wi h GeneMa k-EP + and AUGUSTUS sup- po ed by a p o ein da abase. NAR Genom Bioin o m. 2021;3:lqaa108. 69. S anke M, Schö mann O, Mo gens e n B, Waack S. Gene p edic ion in euka y- o es wi h a gene alized hidden Ma ko model ha uses hin s om ex e nal sou ces. BMC Bioin o ma ics. 2006;7:1–11. 70. S anke M, Diekhans M, Bae sch R, Haussle D. Using na i e and syn enically mapped cDNA alignmen s o imp o e de no o gene inding. Bioin o ma ics. 2008;24:637–44. 71. Ba ne DW, Ga ison EK, Quinlan AR, S ömbe g MP, Ma h GT. BamTools: a C + + API and oolki o analyzing and managing BAM iles. Bioin o ma ics. 2011;27:1691–2. 72. Lomsadze A, Bu ns PD, Bo odo sky M. In eg a ion o mapped RNA-Seq eads in o au oma ic aining o euka yo ic gene inding algo i hm. Nucleic Acids Res. 2014;42:e119–119. 73. Buch ink B, Xie C, Huson DH. Fas and sensi i e p o ein alignmen using DIAMOND. Na Me hods. 2015;12:59–60. 74. Ho KJ, Lange S, Lomsadze A, Bo odo sky M, S anke M. BRAKER1: unsupe - ised RNA-Seq-based genome anno a ion wi h GeneMa k-ET and AUGUS- TUS. Bioin o ma ics. 2016;32:767–9. 75. Shuma e A, Salzbe g SL. Li o : accu a e mapping o gene anno a ions. Bioin o ma ics. 2021;37:1639–43. 76. Li H. Aligning sequence eads, clone sequences and assembly con igs wi h BWA-MEM. a Xi p ep in a Xi :13033997. 2013. 77. Ta aso A, Vilella AJ, Cuppen E, Nijman IJ, P ins P. Sambamba: as p ocessing o NGS alignmen o ma s. Bioin o ma ics. 2015;31:2032–4. 78. Ga ison E, Ma h G. Haplo ype-based a ian de ec ion om sho - ead sequencing. a Xi P ep in a Xi :12073907. 2012. 79. Tan A, Abecasis GR, Kang HM. Uni ied ep esen a ion o gene ic a ian s. Bioin o ma ics. 2015;31:2202–4. 80. Pu cell S, Neale B, Todd-B own K, Thomas L, Fe ei a MAR, Bende D, e al. PLINK: a ool se o whole-genome associa ion and popula ion-based link- age analyses. Am J Hum Gene . 2007;81:559–75. 81. Quinlan AR, Hall IM. BEDTools: a lexible sui e o u ili ies o compa ing genomic ea u es. Bioin o ma ics. 2010;26:841–2. 82. Wickham H, Chang W, Wickham MH. Package ‘ggplo 2.’ C ea e elegan da a isualisa ions using he g amma o g aphics Ve sion. 2016;2:1–189. 83. Va gas-Cha ez C, Pendy NML, Nsango SE, Aguile a L, Ayala D, González J. T ansposable elemen a ian s and hei po en ial adap i e impac in u ban popula ions o he mala ia ec o Anopheles coluzzii. Genome Res. 2022;32:189–202. 84. Ko le R, Gómez-Sánchez D, Schlö e e C. PoPoola ionTE2: compa a i e popula ion genomics o ansposable elemen s using Pool-Seq. Mol Biol E ol. 2016;33:2759–64. 85. Yu T, Huang X, Dou S, Tang X, Luo S, Theu kau WE, e al. A benchma k and an algo i hm o de ec ing ge mline ansposon inse ions and measu ing de no o ansposon inse ion equencies. Nucleic Acids Res. 2021;49:e44–44. 86. G ama es LS, Agapi e J, A ill H, Cal i BR, C osby MA, Dos San os G, e al. FlyBase: a guided ou o highligh ed ea u es. Gene ics. 2022;220:iyac035. 87. Villanue a-Cañas JL, Ho a h V, Aguile a L, González J. Di e se amilies o ansposable elemen s a ec he ansc ip ional egula ion o s ess- esponse genes in D osophila melanogas e . Nucleic Acids Res. 2019;47:6842–57. 88. FlyBase. Da abase [In e ne ]. [ci ed 2022 No 20]. A ailable om: h ps:// ly- base.o g. 89. Salminen TS, Vesala L, Laiho A, Me isalo M, Hoikkala A, Kanka e M. Seasonal gene exp ession kine ics be ween diapause phases in D osophila i ilis g oup species and o e win e ing di e ences be ween diapausing and non- diapausing emales. Sci Rep. 2015;5:1–13. 90. Guo Z, Qin J, Zhou X, Zhang Y. Insec ansc ip ion ac o s: a landscape o hei s uc u es and biological unc ions in D osophila and beyond. In J Mol Sci. 2018;19:3691. 91. Jukam D, Vie s K, Ande son C, Zhou C, DeFo d P, Yan J, e al. The insula o p o ein BEAF-32 is equi ed o Hippo pa hway ac i i y in he e minal di - e en ia ion o neu onal sub ypes. De elopmen . 2016;143:2389–97. 92. Reddy DA, P asad B, Mi a CK. Func ional classi ica ion o ansc ip ion ac o binding si es: in o ma ion con en as a me ic. J In eg Bioin (JIB). 2006;3:32–44. 93. Sandelin A, Alkema W, Engs öm P, Wasse man WW, Lenha d B. JASPAR: an open-access da abase o euka yo ic ansc ip ion ac o binding p o iles. Nucleic Acids Res. 2004;32:D91–4. 94. G an CE, Bailey TL, Noble WS. FIMO: scanning o occu ences o a gi en mo i . Bioin o ma ics. 2011;27:1017–8. 95. Bailey TL, Boden M, Buske FA, F i h M, G an CE, Clemen i L, e al. MEME SUITE: ools o mo i disco e y and sea ching. Nucleic Acids Res. 2009;37:W202–8. 96. Bailey TL, Johnson J, G an CE, Noble WS. The MEME sui e. Nucleic Acids Res. 2015;43:W39–49. 97. Slou skin A, Danino YM, O ens ein Y, Zeha i Y, Donige T, Shami R, e al. Ele- meNT: a compu a ional ool o de ec ing co e p omo e elemen s. T ansc ip- ion. 2015;6:41–50. 98. ElemeNT co e p omo e de ec ing ool [In e ne ]. [ci ed 2022 Sep 11]. h p:// li e acul y.biu.ac.il/ge shon- ama /index.php/ esou ces Page 21 o 21Tahami e al. Mobile DNA (2024) 15:18 99. Hackl T, Ankenb and M, an Ad ichem B. gggenomes: A G amma o G aph- ics o Compa a i e Genomics [In e ne ]. 2024 [ci ed 2024 Jun 20]. h ps:// gi hub.com/ hackl/gggenomes 100. Li H, Du bin R. Fas and accu a e sho ead alignmen wi h Bu ows–Wheele ans o m. Bioin o ma ics. 2010;25:1754–60. 101. Neph S, Kuehn MS, Reynolds AP, Haugen E, Thu man RE, Johnson AK, e al. BEDOPS: high-pe o mance genomic ea u e ope a ions. Bioin o ma ics. 2012;28:1919–20. 102. Pica d pipeline. B oad Ins i u e [In e ne ]. [ci ed 2023 Aug 27]. h p://b oadin- s i u e.gi hub.io/pica d 103. Sedlazeck FJ, Reschenede P, Smolka M, Fang H, Na es ad M, Von Haesele A, e al. Accu a e de ec ion o complex s uc u al a ia ions using single-mole- cule sequencing. Na Me hods. 2018;15:461–8. 104. Rausch T, Zichne T, Schla l A, S ü z AM, Benes V, Ko bel JO. DELLY: s uc u al a ian disco e y by in eg a ed pai ed-end and spli - ead analysis. Bioin o - ma ics. 2012;28:i333–9. 105. Je a es DC, Jolly C, Ho i M, Speed D, Shaw L, Rallis C, e al. T ansien s uc u al a ia ions ha e s ong e ec s on quan i a i e ai s and ep oduc i e isola- ion in ission yeas . Na Commun. 2017;8:14061. 106. Rech GE, Radío S, Gui ao-Rico S, Aguile a L, Ho a h V, G een L, e al. Popula- ion-scale long- ead sequencing unco e s ansposable elemen s associa ed wi h gene exp ession a ia ion and adap i e signa u es in D osophila. Na Commun. 2022;13:1948. 107. Fonseca PM, Mou a RD, Wallau GL, Lo e o ELS. The mobilome o D osophila incomp a, a lowe -b eeding species: compa ison o ansposable ele- men landscapes among gene alis and specialis lies. Ch omosome Res. 2019;27:203–19. 108. Sim C, Denlinge DL. Insulin signaling and FOXO egula e he o e - win e ing diapause o he mosqui o Culex pipiens. P oc Na l Acad Sci. 2008;105:6777–81. 109. Dan o W, Da is MM, Lind all JM, Tang X, U ell H, Junell A, e al. The Oc 1 homolog nubbin is a ep esso o NF-κB-dependen immune gene exp es- sion ha inc eases he ole ance o gu mic obio a. BMC Biol. 2013;11:1–18. 110. Roseman RR, Pi o a V, Geye PK. The su (hw) p o ein insula es exp ession o he D osophila melanogas e whi e gene om ch omosomal posi ion- e ec s. EMBO J. 1993;12:435–42. 111. Tapanainen R, Pa ke DJ, Kanka e M. Pho osensi i e Al e na i e Splicing o he ci cadian clock gene imeless is Popula ion Speci ic in a Cold-adap ed ly, D osophila mon ana. G3 Genes|Genomes|Gene ics. 2018;8:1291–7. 112. Obba d DJ, Maclennan J, Kim K-W, Rambau A, O’G ady PM, Jiggins FM. Es i- ma ing di e gence da es and subs i u ion a es in he D osophila phylogeny. Mol Biol E ol. 2012;29:3459–73. 113. Yusu LH, Tyukmae a V, Hoikkala A, Ri chie MG. Di e gence and in og ession among he i ilis g oup o D osophila. E ol Le . 2022;6:537–51. 114. Lo e o ELS, Ca a e o CMA, Capy P. Re isi ing ho izon al ans e o anspos- able elemen s in D osophila. He edi y (Edinb). 2008;100:545–54. 115. McCulle s TJ, S einige M. T ansposable elemen s in D osophila. Mob Gene Elem. 2017;7:1–18. 116. Rius N, Guillén Y, Delp a A, Kapus a A, Fescho e C, Ruiz A. Explo a ion o he D osophila buzza ii ansposable elemen con en sugges s unde es ima ion o epea s in D osophila genomes. BMC Genomics. 2016;17:1–14. 117. Maumus F, Fis on-La ie A-S, Quesne ille H. Impac o ansposable elemen s on insec genomes and biology. Cu Opin Insec Sci. 2015;30–6. 118. Ellison CE, Bach og D. Dosage compensa ion ia ansposable elemen media ed ewi ing o a egula o y ne wo k. Science (1979). 2013;342:846–50. 119. Viei a C, Na don C, A pin C, Lepe i D, Biémon C. E olu ion o genome size in D osophila. Is he in ade ’s genome being in aded by ansposable ele- men s? Mol Biol E ol. 2002;19:1154–61. 120. S i C, Go don SP, Wicke T, Vogel JP, Roulin AC. Recen ac i i y in expanding popula ions and pu i ying selec ion ha e shaped ansposable elemen land- scapes ac oss na u al accessions o he Medi e anean g ass B achypodium dis achyon. Genome Biol E ol. 2018;10:304–18. 121. Me enciano M, Iacome i C, González J. A unique clus e o oo inse ions in he p omo e egion o a s ess esponse gene in D osophila melanogas e . Mob DNA. 2019;10:1–11. 122. Psenako a K, Kohou o a K, Obsilo a V, Ausse lechne MJ, Ve e ka V, Obsil T. Fo khead domains o FOXO ansc ip ion ac o s di e in bo h o e all con o - ma ion and dynamics. Cells. 2019;8:966. 123. Kapun M, Fla T. The adap i e signi icance o ch omosomal in e sion poly- mo phisms in D osophila melanogas e . Mol Ecol. 2019;28:1263–82. 124. Poikela N, Lae sch DR, Kanka e M, Hoikkala A, Lohse K. Expe imen al in o- g ession in D osophila: asymme ic pos zygo ic isola ion associa ed wi h ch omosomal in e sions and an incompa ibili y locus on he X ch omosome. Mol Ecol. 2023;32:854–66. Publishe ’s no e Sp inge Na u e emains neu al wi h ega d o ju isdic ional claims in published maps and ins i u ional a ilia ions.