scieee Open visual document viewer

G-quadruplexes in the evolution of hepatitis B virus

Brázda, Václav; Dobrovolná, Michaela; Bohálová, Natália; Mergny, Jean-Louis

Abstract

Hepatitis B virus (HBV) is one of the most dangerous human pathogenic viruses found in all corners of the world. Recent sequencing of ancient HBV viruses revealed that these viruses have accompanied humanity for several millenia. As G-quadruplexes are considered to be potential therapeutic targets in virology, we examined G-quadruplex-forming sequences (PQS) in modern and ancient HBV genomes. Our analyses showed the presence of PQS in all 232 tested HBV genomes, with a total number of 1258 motifs and an average frequency of 1.69 PQS per kbp. Notably, the PQS with the highest G4Hunter score in the reference genome is the most highly conserved. Interestingly, the density of PQS motifs is lower in ancient HBV genomes than in their modern counterparts (1.5 and 1.9/kb, respectively). This modern frequency of 1.90 is very close to the PQS frequency of the human genome (1.93) using identical parameters. This indicates that the PQS content in HBV increased over time to become closer to the PQS frequency in the human genome. No statistically significant differences were found between PQS densities in HBV lineages found in different continents. These results, which constitute the first paleogenomics analysis of G4 propensity, are in agreement with our hypothesis that, for viruses causing chronic infections, their PQS frequencies tend to converge evolutionarily with those of their hosts, as a kind of 'genetic camouflage' to both hijack host cell transcriptional regulatory systems and to avoid recognition as foreign material.

Full text

7198–7204 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 Published online 3 July 2023 h ps://doi.o g/10.1093/na /gkad556 G-quad uplexes in he e olu ion o hepa i is B i us V ´ acla B ´ azda 1 , * , Michaela Dob o oln ´ a 1 , 2 , Na ´ alia Boh ´ alo ´ a 1 and Jean-Louis Me gny 1 , 3 , * 1 Ins i u e o Biophysics o he Czech Academy o Sciences, B no, Czech Republic, 2 Facul y o Chemis y, B no Uni e si y o Technology, Pu ky ˇ no a 118, 612 00 B no, Czech Republic and 3 Labo a oi e d’Op ique e Biosciences (LOB), Ecole Poly echnique, CNRS, INSERM, Ins i u Poly echnique de Pa is, 91120 Palaiseau, F ance Recei ed Ma ch 30, 2023; Re ised May 23, 2023; Edi o ial Decision June 14, 2023; Accep ed June 19, 2023 ABSTRACT Hepa i is B i us (HBV) is one o he mos dange ous human pa hogenic i uses ound in all co ne s o he wo ld. Recen sequencing o ancien HBV i uses e- ealed ha hese i uses ha e accompanied human- i y o se e al millenia. As G-quad uplexes a e con- side ed o be po en ial he apeu ic a ge s in i ol- ogy, we examined G-quad uplex- o ming sequences (PQS) in mode n and ancien HBV genomes. Ou anal yses sho wed he p esence o PQS in all 232 es ed HBV genomes, wi h a o al numbe o 1258 mo- i s and an a e a ge equenc y o 1.69 PQS pe kbp. No ably, he PQS wi h he highes G4Hun e sco e in he e e ence genome is he mos highly conse ed. In e es ingly, he densi y o PQS mo i s is lowe in ancien HBV genomes han in hei mode n coun e - pa s (1.5 and 1.9 / kb, espec i ely). This mode n e- quency o 1.90 is e y close o he PQS equency o he human genome (1.93) using iden ical pa ame e s. This indica es ha he PQS con en in HBV inc eased o e ime o become close o he PQS equency in he human genome. No s a is ically signi ican di - e ences we e ound be ween PQS densi ies in HBV lineages ound in di e en con inen s. These esul s, which cons i u e he i s paleogenomics analysis o G4 p opensi y, a e in ag eemen wi h ou hypo he- sis ha , o i uses causing ch onic in ec ions, hei PQS equencies end o con e ge e olu iona il y wi h hose o hei hos s, as a kind o ‘gene ic camou lage’ o bo h hijack hos cell ansc ip ional egula o y sys- ems and o a oid ecogni ion as o eign ma e ial. GRAPHICAL ABSTRACT INTRODUCTION Hepa i is B i us (HBV) belongs o he genus O hohep- adna i us; i s genome is cons i u ed o double-s anded DNA. This i us causes Hepa i is B , a highly con a- gious, po en ially a al disease ha a ec s an es ima ed 257 million people wo ldwide, esul ing in an es ima ed 820 000 dea hs e e y yea ( h ps://www.who.in /news- oom/ ac - shee s/de ail/hepa i is- b ). This i us is pa icula ly se- e e, as a pp oxima el y one in i e ca ie s die om ci ho- sis and / o de elop hepa ocellula ca cinoma. HBV is ans- mi ed p ima ily h ough blood and body luids and he incuba ion pe iod is a iable, usually be ween 30 and 180 days. Du ing eplica ion, HBV DNA o ms a minich omo- some in he nucleus o in ec ed hepa ocy es ( 1 , 2 ) and i s genome is eplica ed h ough a p ocess o e e se ansc ip- ion o he key in e media e p e-genomic RNA in hepa o- cy es, which is also an mRNA empla e o he HBV p o- eins ( 3 , 4 ). Hepadna i uses in ec ing o he hos s ha e e- cen ly been iden i ied, including ba s , ogs , liza ds , ish, and he capuchin monkey ( 5–8 ). Analyses o ancien genomes ha e e ealed ha he mos ecen common ances o o all HBV lineages is es ima ed o ha e exis ed be ween ∼20 000 * To whom co espondence should be add essed. Tel: +420 541517231; Email: acla @ibp.cz Co espondence may also be add essed o Jean-Louis Me gny. Tel: +33 766290967; Email: jean-louis.me gny@poly echnique.edu C The Au ho (s) 2023. Published by Ox o d Uni e si y P ess on behal o Nucleic Acids Resea ch. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p: // c ea i ecommons.o g / licenses / by / 4.0 / ), which pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7199 and 12 000 yea s ago, and he i us was ound o be p esen in Eu opean and Sou h Ame ican hun e -ga he e s du ing he ea ly Holocene pe iod ( 9 ). A la ge amoun o li e a u e is de o ed o he occu ence o G-quad uplexes in i uses, especially o he possibili y o using hese s uc u es in he apy. Comp ehensi e bioin o - ma ics analyses ha e aced pu a i e G4- o ming sequences in he genome o almos all human i uses, showing ha hei dis ibu ion and p esence a e highly conse ed. The e- o e, hese DNA o RNA s uc u es can be a sui able a - ge o a ge ed he apy. Some G-quad uplex ligands ha e been shown o ha e an i i al ac i i y, o example, agains HIV ( 10 ), he pes simplex i us I (HSV-1) ( 11 ), SARS-CoV- 2 ( 12 ) and o he s. A comp ehensi e analysis o all sequenced i uses ha ha e a la en phase in hei li e cycle showed ha hei G-quad uplex con en is co ela ed wi h ha o he hos ( 13 ). In con as , i uses causing acu e in ec ions wi hou a la en phase end o elimina e G-quad uplex se- quences, as hey can become oadblocks du ing eplica ion, ansc ip ion and / o e e se ansc ip ion ( 14 ). Mo e speci ically, G4s ha e been ound o be ele an in HBV in ec ion, bo h a he DN A and RN A le el ( 15 ). Chak abo y and Ghosh epo ed ha an RNA sequence p esen in HBV RNA exhibi ed a sequence-independen ans-ac ing nuclease ac i i y, and ha his sequence adop s a G4 con o ma ion ( 16 ). Biswas e al. analyzed a G4 p one mo i ( GGGAGTGGGAGCATTCGGGCCAGGG ) ha is highly conse ed only in HBV geno ype B, and was shown o adop a hyb id s uc u e ( 17 ). In e es ingl y, m u a ions dis- up ing his G-quad uplex in HBV geno ype B cons uc s we e associa ed wi h impai ed i ion sec e ion. The au ho s p oposed ha his G4 media es enhancemen o ansc ip- ion and i ion sec e ion in his HBV geno ype. In a la e e- iew, hey no e ha among i uses con aining a G4 in hei genome, hose associa ed wi h cance a e o e - ep esen ed, including HBV ( 18 ). In con as , some G4s end o be conse ed in all geno- ypes, as epo ed by Meie -S ephenson e al ( 19 ) o a DNA sequence ound in he p e-co e p omo e egion ( CTGGGAGGAGCTGGGGGAGGAGA ). They demons a ed a ole o his quad uplex in i al eplica ion by compa ing he wild- ype mo i o non-G4- o ming mu an s in i o . Fleming e al iden i ied conse ed po en ial G4 sequences in se e al i al genomes ele an o human heal h, and showed ha hese mo i s can p o ide a ame wo k o N6- me hyladenosine (m6A) ins alla ion wi hin he loops o RNA G4 sequences ound in se e al i uses, including HBV ( 20 ). Somku i e al. de e mined olume changes in h ee HBV G4 s uc u es using biophysical app oaches in i o . They in es iga ed h ee DNA sequences: GGCTGGGGCTTG- GTCATGGGCCATCAG , GGGAGTGGGAGCATTCGGGCCAGGG and TTGGGTGGCTTTGGGGCATGGAC ( 21 ). The same g oup in es iga ed one o hese sequences in mo e de ails (HepB; GGCTGGGGCTTGGTCATGGGCCATCAG , ound in he coding egion o he polyme ase p o ein), and analyzed i s in e - ac ion wi h G4 ligands in i o ( 22 ). Finally, Sun e al. e- cen ly in es iga ed he ole o cellula G4s (mo i s ound in he hos cell genome) in HBV in ec ion, demons a ing ha he DDX5 helicase, known o be capable o esol ing RNA G4 s uc u es, is a key egula o o he in e e on (IFN) esponse agains his i us. DDX5 down egula ion is ob- se ed du ing HBV eplica ion and in poo p ognosis HBV- ela ed hepa ocellula ca cinoma (HCC) ( 23 ). All o hese esul s poin ou he links be ween i al (o hos ) DNA and RNA G4s and HBV. In his pape , we ha e analyzed 232 HBV genomes om samples co e ing a mo e han 10- housand-yea his o y o he p esence o G-quad uplex o ming sequences. Ou e- sul s show an e olu iona y shi o an inc eased numbe o G-quad uplex es in ecen HBV i uses, poin ing o he im- po ance o G-quad uplexes in he HBV li e-cycle in human li e cells. MATERIALS AND METHODS Genomes 232 HBV alignmen s we e downloaded om he supple- men a y ma e ials a (9). Sequences we e ob ained o 122 mode n geno ypes and di e en g oups o 110 ancien s ains, di ided in o g oups based on mPTP classi ica ion. As he e e ence genome, we ook NC 003977.2 and, o analyses o phylogene ically ela ed i uses wi h hos s o he han human, we il e ed e e ence genomes om O hohep- adna i uses. In o al, we downloaded 21 addi ional HBVs ha ing a non-human hos , in ec ing bi ds (7 genomes), ba s (5 genomes), ish (2 genomes), o he Mammals (i.e. nei he human no ba s; 6 genomes) and amphibians (1 genome, Ti- be an og hepa i is B i us). HBV G4 con en s we e com- pa ed o he gapless human genome, he new elome e- o- elome e assembly o he human genome ( 24 ), which was downloaded om NCBI (T2T-CHM13 2.0). G4Hun e analyses All sequences we e analyzed using G4Hun e ( h p://bioin o ma ics .ibp .cz ) o iden i y PQS sequences. G4Hun e ’s de aul pa ame e s we e used (25 nucleo ides o window size and 1.2 o h eshold). These se ings ha e p e iously been shown o iden i y expe imen ally- alida ed quad uplex s uc u es. The lis o all o ganisms es ed and he esul s o he analyses we e downloaded om he supplemen a y ma e ials a ( 9 ). S a is ical e alua ion Da a wi h G4Hun e esul s we e me ged in an Excel ile o s a is ical e alua ion. G4Hun e esul s , leng hs , and GC con en o analyzed sequences a e accessible in Supplemen- a y ma e ial 01. A sca e plo was gene a ed in G aph- Pad P ism ( 8.0.1), Violin plo s we e cons uc ed in R ( 4.2.0) wi h ggplo 2. S a is ical signi icance was es ed using S uden ’s T- es . No mali y o da a was de e mined using Shapi o-Wilk es . Cons uc ion o LOGO sequence All sequences o ancien and mode n HBV genomes we e uploaded in o UGENE so wa e ( 25 ) and he loca ion o PQS sequences we e ex ac ed using Clus alW alignmen . LOGO sequence was gene a ed in aligned sequences and WebLogo 3 ool ( 26 ). Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024 7200 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 Table 1. S a is ic da a o G4Hun e analyses o HBV i uses. n Seq (numbe o s ains), Leng h (leng h o he sequence, n ), GC % (a e age GC con en ), PQS n ( o al numbe o p edic ed PQS wi h a G4Hun e sco e o 1.2 o mo e), Mean PQS (a e age numbe o p edic ed PQS), Min PQS (lowes equency o p edic ed PQS), Max PQS (highes equency o p edic ed PQS), PQS pe 1000 GC (PQS equency pe 1000 GC) G oups Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC All HBV 232 3214 45.8 1258 1.69 0.61 4.04 3.66 G oups Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC ancien 122 3217 43.1 587 1.50 0.61 2.83 3.46 mode n 110 3210 48.8 671 1.90 0.94 4.04 3.88 H. sapiens T2T 1* 3.05 ×10 9 41.6 5.45 10 6 1.93 - - 4.72 Subg oups** Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC A 21 3227 47.0 105 1.55 2.00 8.00 3.30 B 14 3222 49.0 120 2.66 1.55 4.04 5.43 C 16 3222 49.1 96 1.86 1.55 3.10 3.79 D 57 3190 47.7 317 1.74 0.94 2.83 3.66 E 3 3218 48.4 25 2.59 2.17 2.80 5.35 F 13 3222 48.8 81 1.93 1.24 2.48 3.96 G 3 3255 48.2 12 1.23 1.23 1.23 2.55 H 5 3222 48.9 39 2.42 1.86 3.11 4.94 I 3 3219 48.8 20 2.07 1.55 2.49 4.24 J 1 3187 49.0 4 1.26 1.26 1.26 2.56 Ancien Ame ican 4 3214 38.4 15 1.17 0.93 1.55 3.12 Ea ly Ana olian a me 1 3189 33.7 2 0.63 0.63 0.63 1.86 Mesoli hic 8 3192 45.8 33 1.29 0.63 2.20 3.20 WENBA 65 3229 42.1 288 1.37 0.61 2.20 3.27 o u 2 3191 49.1 8 1.25 1.25 1.25 2.56 gbn 3 3194 49.1 20 2.09 1.57 2.51 4.25 czp 3 3191 48.3 18 1.88 1.25 2.51 3.89 o he 10 3211 45.2 55 1.71 0.62 4.04 3.81 *Comple e elome e- o- elome e human genome (22 + X + Y ch omosomes) **The g oups we e de ined by Koche e al. ( 10 ) and is based on he mul i- a e Poisson T ee P ocesses (mPTP) as a esul s o gene ic clus e s numbe conside ing a phylogene ic inpu ee ( 26 ). RESULTS We analyzed he p esence o PQS in 232 HBV genomes (122 ancien and 110 ecen ) using G4Hun e . All human HBV genomes a e simila in leng h, a ying om 3180 o 3300 bp. Compa isons o ancien and ecen samples show a sligh and non-signi ican change in a e age leng h om 3217 o 3210 bp. On he o he hand, hese genomes a y in GC con en , and we ound he p esence o G-quad uplex- o ming sequences in all HBV genomes in he da ase . In o al, we ound 1258 PQS wi h a mean equency o 1.69 PQS pe kbp (Table 1 ). The mean equency o PQS in ancien genomes is 1.50 / kb, compa ed wi h he mean e- quency in mode n HBV genomes o 1.90. Fo compa ison we also analyzed he ne wly pub lished human gapless assem- bly. The PQS equency in he human genome is 1.93, which is almos iden ical o he a e age PQS equency o mode n HBV genomes (Table 1 ). Al hough he e is no signi ican change in he leng h o ancien and mode n HBV genomes, compa ison o PQS densi y shows ha he mode n i uses a e subs an ially iche in PQS (Figu e 1 ) ( P - alue = 3.6e- 08). The mode n HBV genomes no only ha e a highe PQS equency, bu also ha e a highe GC con en . To e alua e i he change in PQS equency is s a is ically signi ican a e aking in o accoun GC con en , we ecalcula ed he PQS equency acco ding o GC con en (Table 1 , las column). E en a e his co ec ion, he PQS equency / GC con en is highe in mode n han in ancien HBV genomes (3.88 e - sus 3.46 pe housand GC o he mode n and ancien HBV genomes, espec i ely: P - alue = 1.8e-03). We hen u he di ided genomes acco ding o di e - en geno ypes. While mos cu en geno ypes ha e an a - Figu e 1. Compa ison o PQS equencies in ancien and mode n HBV genomes. e age PQS equency highe han 2 PQS / kb, ou o he i e ancien genomes ha e PQS equencies lowe han his alue. The highes PQS equencies we e ound in mod- e n geno ypes B, E and H, he lowes in ancien Ame ican, Mesoli hic and o he ancien geno ypes (Figu e 2 ). We p esen he PQS equency pe kb ( o all mo i s wi h a G4Hun e sco e > 1.2) and he leng h o he genome as a unc ion o ime o each ancien sequence (Supplemen a y ma e ial 02). As men ioned be o e, he longes sequence only di e s om he sho es by 120 bp, and he leng h o he genome does no change signi ican ly o e ime (Pea - son = −0.1387, P ( wo- ailed) = 1.6e-01). Unlike genome leng h, PQS equencies pe kbp we e ound o inc ease o e Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7201 Figu e 2. PQS equency o HBV subg oups di ided in o ancien and mode n geno ypes. Ancien Ame ican 1, 2, 3 ca ego ies we e me ged. Mesoli hic 1 and 2 geno ypes we e me ged. Table 2. Posi ion o G4 sequences in he human HBV e e ence genome (NC 003977.2) and i s loca ion conse a ion. A nega i e sco e co esponds o a C- ich sequence, meaning ha i is i s complemen a y s and which will be G4-p one Posi ion S and Sequence G4H sco e Fea u e / loca ion G4 conse a ion 316 −CCCC AA CC T CC AAT C A C T C A CC −1.41 S, DNA polyme ase N- e minal domain 84.5% 737 + GG AT G AT G T GG TATT GGGGG +1.50 S, DNA polyme ase C- e minal domain 70.3% 1133 −CC TGAA CC TTTA CCCC GTTG CCC −1.30 P, DNA polyme ase C- e minal domain 19.4% 1722 + GGG A GG AGTT GGGGG A GG A G ATTA GG +1.65 T ansac i a ion p o ein X 92.2% 1887 + GGG T GG CTTT GGGG CAT GG +1.63 C , Hepa i is co e p o ein 40.5% ime (Pea son = 0.2903, P ( wo- ailed) = 2.8e-03; Supple- men a y ma e ial 02). The whole genome o HBV is ansc ibed in o one long p e-genomic RN A (mRN A-pgRN A) ha encodes all HBV p o eins, and pgRNA also ac s as a empla e o e e se ansc ip ion. Analysis o PQS localiza ion in he HBV e - e ence genome (NC 003977.2) shows ha all PQS a e lo- ca ed in he icini y o gene egions, which is no su p ising conside ing ha he HBV genome is small and he en i e genome is used e y e ec i ely o p oduce he ew p o eins necessa y o i s unc ion, such as DN A pol yme ase, ans- ac i a ion and capsid p o eins (Table 2 ). The PQS wi h he highes G4Hun e sco e in he HBV e e ence genome is loca ed su ounding posi ion 1722, in he egion ha codes o ansac i a ion p o ein X. In e - es ingly, a PQS is p esen in almos all HBV genomes a his loca ion (214 o 232; o 92.2% o HBV genomes an- alyzed), making i he mos conse ed mo i be ween all PQS in he e e ence HBV genome. Conse a ion implies ha his sequence posi ion has been main ained by selec- i e p essu e. Compa ison o he LOGO sequence o his loca ion in mode n and ancien HBV genomes (Figu e 3 ) demons a es ha a G- ich mo i is p ese ed in all s ains. Ne e heless, his ‘G- ichness’ is e en mo e s iking in mod- e n compa ed o ancien s ains, wi h Gs becoming p edom- inan a posi ions 8 and 11 (Figu e 3 , a ows), while ancien Figu e 3. L OGO ep esen a ion o he consensus mo i ound a ound po- si ion 1722 (acco ding o he e e ence HBV genome) in mode n and an- cien HBV genomes. Guanine nucleo ides p edomina e in mode n s ains a posi ions 8 and 11 (a ows) while in ancien HBV genomes o he nu- cleo ides a e p esen (T and A in posi ion 8). A less s iking G-en ichmen is also ound a posi ions 14 and 15. As a consequence, while bo h con- sensus mo i s a e compa ible wi h G4 o ma ion, he mode n sequence is mo e a o able han he ancien one (G4hun e sco es o 2.11 and 1.83, espec i ely). Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024 7202 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 Table 3. G4Hun e sco e in HBV i uses, g ouped by con inen and age G oups Seq Leng h Mean PQS Min PQS Max PQS All 232 3214 1.69 0.61 4.04 Ancien Seq Leng h Mean PQS Min PQS Max PQS Aus alia 0 - - - - Ame ica 7 3217 1.33 0.93 3.11 A ica 1 3229 0.62 - - Asia 34 3219 1.54 0.61 2.51 Eu ope 82 3217 1.50 0.61 2.83 Mode n Seq Leng h Mean PQS Min PQS Max PQS Aus alia 14 3215 1.93 0.94 3.10 Ame ica 31 3216 1.96 0.94 3.11 A ica 15 3205 1.83 1.24 2.80 Asia 110 3206 1.93 0.94 4.04 Eu ope 12 3209 1.66 1.23 2.48 HBV genomes mo e o en exhibi T / A o A a hese posi- ions. A highe –– nea 100% –– p e alence o Gs is also is- ible a posi ions 14 and 15 (Figu e 3 ) in mode n genomes. O e all, while bo h ancien and mode n consensus mo i s a e G4-p one, he mode n sequence is mo e a o able, as shown by he highe G4Hun e sco e. We also di ided HBV genomes acco ding o he geo- g aphic place o sampling. Mos o he ancien HBV sam- ples we e ound in Eu ope, only one in A ica and none in Aus alia. None heless, ancien genomes ha e a lowe PQS equency han mode n HBV genomes, ega dless o hei con inen o o igin (Table 3 ; also see G aphical Abs ac ). DISCUSSION Pa hogens e ol e in esponse o human biological changes alongside sociocul u al and echnological de elopmen s ( 27 ). Ancien i al genomes p o ide in o ma ion on he e olu ion o i uses o e bo h ime and space and p o- ide insigh in o he changes ha may ha e occu ed in i ulence and ansmissibili y ( 28 ). Cu en ad anced ech- niques o isola ion o nucleic acids and sequencing ha e al- lowed paleogenomic o ‘a cheo i ology’ in es iga ions. The in amous 1918 ‘Spanish’ in luenza pandemic was he sou ce o he i s ancien pa hogen genome ( 29 ) and a cheo i ol- ogy has been g owing apidly since hen. Conside ing ha wo- hi ds o all human pa hogens a e i uses ( 30 ), paleo- gene ics o i al genomes p o ides an in e es ing iewpoin on human his o y ( 31 ). While se e al ancien i al genomes a e a ailable, he mos comp ehensi e da ase deals wi h ancien HBV genomes ( 28 ). Wi h mo e han 200 million people su e ing om ch onic HBV in ec ion, HBV can be conside ed as a common i us. HBV has a li e-cycle ha equi es i s double- s anded DNA genome o each he hos cell nucleus ( 32 ), in con as wi h RNA i uses such as in luenza o SARS- CoV-2 ha cause only acu e in ec ions and con ain RNA genomes ha can be eplica ed and ansla ed in he cy o- plasm. Vi uses causing acu e in ec ions end o ha e a low PQS equency, while G4s end o be ubiqui ous in mos o - ganisms. G4s in he HBV genome a e impo an s uc u es ha egula e ansc ip ion and i ion sec e ion in HBV geno- ype B ( 15 , 17 ). In his epo , we analyzed he p esence o PQS in mul iple HBV genomes, om ancien o cu en HBV s ains. Impo an ly, we ound ha PQS equency is highe in ecen compa ed o ancien s ains. I was shown p e iously ha he G4 equency o dsDNA i uses co e- la es wi h he PQS equency o he hos , shown o dsDNA i uses in ec ing A chaea, Bac e ia and Euka yo a ( 33 ). In ag eemen wi h hose da a, we ound ha he densi y o G4 mo i s in mode n HBV s ains (dsDNA i uses ha also expe ience a la en phase) ends o con e ge o he o e all G4 densi y o he human genome. We p opose ha mu a- ions which led o ‘PQS in eg a ion’ (new G4 mo i s wi hin he HBV genome) we e e olu iona y p e e ed and ixed in HBV pa hogenic s ains du ing e olu ion. I seems ha he opposi e p ocess may occu in i uses causing acu e in ec- ions, as ound o SARS-CoV-2, whe e he PQS equency is ex emely low compa ed o he PQS equencies o o he co ona i uses ( 34 , 35 ). Vi uses ha e highly a iable genomes and a e p one o mu a ions, in con as o cellula and especially mul icellu- la o ganisms, as desc ibed epea edly. This is especially ue o RN A i uses, w he e he m u a ion a e is se e al o de s o magni ude highe han in DNA-based genomes. A d a- ma ic dec ease in GC con en has been desc ibed o se e al bac e ial species ( 36 ) and o some plan species wi h holo- cen ic ch omosomes ( 37 ). Simila ly, a apid dec ease in GC con en o e ime was ound o some i uses causing acu e in ec ions (and wi hou la ency connec ed o nuclea local- iza ion), including Nido i ales, in luenza genomes and he con empo a y SARS-CoV-2 ou b eak ( 38 ). Compa ison o SARS-CoV-2 genomes showed a s ong p e e ence o mu- a ions in GC islands and C > U ansi ions, leading o a dec ease in G4 p opensi y ( 39 , 40 ). As a consequence, he PQS equency o hese i uses causing only acu e in ec- ions is gene ally e y low (0.03 o SARS-2 ( 35 ), 0.56 o in luenza H1N1 genomes ( 41 )). The opposi e end is ue o HBV, which exhibi s a la- en s a e and main ains i s genome in he nucleus: mode n HBV genomes ha e a signi ican ly highe PQS equency compa ed o ancien HBV genomes. Ou esul s a e in line wi h a b oad s udy compa ing PQS equencies in i uses wi h p edominan ly pe sis en o acu e ypes o in ec ion ( 42 ). Acco ding o ha s udy, i uses causing pe sis en in- ec ion a e en iched in PQS compa ed o acu ely in ec ious i uses. Impo an ly, his obse a ion is also alid wi hin i uses causing hepa i is: HAV (hepa i is A i us - caus- ing acu e in ec ion) ha e a low PQS equency, while he Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7203 Hepadna i idae ha cause ch onic in ec ions ( o which hu- man HBV belongs) ha e a signi ican ly highe PQS e- quency ( 42 ). The p esence o G4s depends on guanine con en in he genome, and one o he possible ad an ages o a GC- ich genome is o p o ide addi ional gene egula ion oppo u- ni ies. In his espec , G4s ha e been shown o be impo - an o ansc ip ion in highe o ganisms. An inc ease in GC con en has also been documen ed in plan species ha can g ow in seasonally cold clima es, possibly indica ing an ad an age o GC- ich DNA du ing cell eezing, and he genomic adap a ions associa ed wi h changing GC con- en a e sugges ed o g ass-domina ed biomes du ing he Te ia y pe iod ( 37 ). G4s a e o en o e ep esen ed in he p omo e egions o highe euka yo es and ha e also been demons a ed o con ibu e o di ec ed genome edi ing in nema odes. Fo i uses expe iencing a la en phase in he nucleus, ha ing a simila genome o ganiza ion as he hos is ad an ageous o bo h a oid ecogni ion as unusual ( o - eign) DNA, and o hijack he hos egula o y machine y. In addi ion, he high GC con en o HSV DNA is sugges ed o ac as a p o ec i e ea u e agains e o ansposon inse ion ( 43 ). As AT- ich egions in humans a e mos ly associa ed wi h condensed ch oma in ( 44 ), he shi o GC- ich i uses could be impo an o i uses wi h la en phase o ha e a be e chance o being ac i e in he u u e. HBV has a la- en pe iod, he e o e, he e olu iona y p essu e o inc ease GC con en and PQS p esence could be e olu iona y a- o ed. In his model, he o iginal (non-human) p e-HBV hos could ha e had a lowe G4 equency – adap a ion o he human hos may ha e led o a hos -pa hogen PQS con- e gence and a concomi an inc ease in G4 densi y in he i us. CONCLUSION We pe o med he i s paleogenomic analysis o G4 p opensi y, applied he e o he Hepa i is B i us. We ound ha he densi y o PQS mo i s inc eased o e ime, as i is highe in mode n han ancien HBV genomes. The e- quency in mode n i uses is now e y close o ha o he human genome. This s udy should pa e he way o he paleo genomics anal ysis o G4 sequences (and o he sim- ila mo i s) in o he pa hogens. Un o una ely, his o ical in o ma ion abou i uses genomes is only a ely a ail- able: A cheo i ology is a nascen ield ( 45 ), which aces he same obs acles as mode n genomics, bu wi h he ad- di ional p oblem o analyzing pa ially deg aded DNA. As no ed by he au ho s ( 45 ), HBV is ‘an excellen a ge o he eco e y o ancien sequences due o i s ela i ely s ab le, pa ially double-s anded ci cula DNA genome, i s high p e alence in he human popula ion, and p olonged high i emia du ing ch onic in ec ion’. Double-s anded DNA is gene ally be e p ese ed han single-s anded DNA o RNA. This may explain why HBV is cu en ly he only in- s ance in which sequence da a a e a ailable o o e 100 an- cien i uses. In o he si ua ions, such as a iola i us, only a ew ancien genomes a e a ailable ( he e a e only 4, 10 and 11 HSV-1, Pa o i uses and VA RV ancien sequences a ailab le, espec i el y ( 28 ). Anal yses o HIV-1 o he 1918 in luenza i uses is possible, bu only o e a limi ed pe iod o ime. A cheo i ology will be use ul o iden i y polymo phisms impo an o human adap a ion o pa hogens, and ice- e sa du ing he complex i us-hos ela ionships c ucial o con inued i al p e alence ( 28 ). Mo e paleogenomic da a will be needed o es he hypo hesis ha , o i uses causing ch onic in ec ions, hei PQS equencies end o con e ge e olu iona ily wi h hose o hei hos . In pa icula , gi en ecen se ious ou b eaks, we hope ha he analysis o an- cien i al pa hogens will p o ide c i ical knowledge abou he na u e o new i al diseases. DA T A A V AILABILITY All da a a e a ailable in he manusc ip and supplemen a y iles. HBV sequences we e uploaded om he supplemen- a y ma e ials a ( 9 ). SUPPLEMENT ARY DA T A Supplemen a y Da a a e a ailable a NAR Online. ACKNOWLEDGEMENTS The au ho s hank P. Coa es o English p oo eading and aluable commen s, and L. Gui a (LOB) o help ul dis- cussions. FUNDING Agence de l’Inno a ion de D ´ e ense (AID) ia he Cen e In e disciplinai e d’E udes pou la D ´ e ense e la S ´ ecu i ´ e (CIEDS) [p ojec 2023 - Pa hogens]; ANR G4Access [ANR-20-CE12-0023]; INCa G4Access g an s ( o J.L.M.); SYMBIT p ojec [CZ.02.1.01 / 0.0 / 0.0 / 15 003 / 0000477] i- nanced om he ERDF; Czech Science Founda ion [22- 21903S o V.B.]. Funding o open access cha ge: Academ y o Sciences. Con lic o in e es s a emen . None decla ed. REFERENCES 1. Bock,C.T., Sch anz,P., Sch ¨ ode ,C.H. and Zen g a ,H. (1994) Hepa i is B i us genome is o ganized in o nucleosomes in he nucleus o he in ec ed cell. Vi us Genes , 8 , 215–229. 2. Newbold,J.E., Xin,H., Tencza,M., She man,G., Dean,J., Bo w den,S. and Loca nini,S. (1995) The co alen ly closed duplex o m o he hepadna i us genome exis s in si u as a he e ogeneous popula ion o i al minich omosomes. J. V i ol. , 69 , 3350–3357. 3. Summe s,J. and Mason,W.S. (1982) Replica ion o he genome o a hepa i is B–like i us by e e se ansc ip ion o an RNA in e media e. Cell , 29 , 403–415. 4. Tiollais,P., Pou cel,C. and Dejean,A. (1985) The hepa i is B i us. Na u e , 317 , 489–495. 5. MacDonald,D.M., Holmes,E.C., Lewis,J.C. and Simmonds,P. (2000) De ec ion o hepa i is B i us in ec ion in wild-bo n chimpanzees (Pan oglody es e us): phylogene ic ela ionships wi h human and o he p ima e geno ypes. J. Vi ol. , 74 , 4253–4257. 6. D exle ,J.F., Geipel,A., K ¨ onig,A., Co man,V.M., an Riel,D., Leij en,L.M., B eme ,C.M., Rasche,A., Co on ail,V.M., Maganga,G.D. e al. (2013) Ba s ca y pa hogenic hepadna i uses an igenically ela ed o hepa i is B i us and capable o in ec ing human hepa ocy es. P oc. Na l. Acad. Sci. U.S.A. , 110 , 16151–16156. Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024 7204 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7. Laube ,C., Sei z,S., Ma ei,S., Suh,A., Beck,J., He s ein,J., B ¨ o old,J., Salzbu ge ,W., Kade ali,L., B iggs,J.A.G. e al. (2017) Deciphe ing he o igin and e olu ion o hepa i is B i uses by means o a amily o non-en eloped ish i uses. Cell Hos Mic obe , 22 , 387–399. 8. de Ca alho Dominguez Souza,B.F., K ¨ onig,A., Rasche,A., de Oli ei a Ca nei o,I., S ephan,N., Co man,V.M., Roppe ,P.L., Goldmann,N., Keppe ,R., M¨ulle ,S.F. e al. (2018) A no el hepa i is B i us species disco e ed in capuchin monkeys sheds new ligh on he e olu ion o p ima e hepadna i uses. J. Hepa ol. , 68 , 1114–1122. 9. Koche ,A., Papac,L., Ba que a,R., Key,F.M., Spy ou,M.A., H¨uble ,R., Roh lach,A.B., A on,F., S ahl,R., Wissgo ,A. e al. (2021) Ten millennia o hepa i is B i us e olu ion. Science , 374 , 182–188. 10. Pe one,R., Bu o skaya,E., Daelemans,D., Pal`u,G., Pannecouque,C. and Rich e ,S.N. (2014) An i-HIV-1 ac i i y o he G-quad uplex ligand BRACO-19. J. An imic ob. Chemo he . , 69 , 3248–3258. 11. F asson,I., Sold ` a,P., Nadai,M., Tassina i,M., Scalab in,M., Gokhale,V., Hu ley,L.H. and Rich e ,S.N. (2022) Quindoline-de i a i es display po en G-quad uplex-media ed an i i al ac i i y agains he pes simple x i us 1. An i i al R es. , 208 , 105432. 12. Zhai,L.-Y., Su,A.-M., Liu,J.-F., Zhao,J.-J., Xi,X.-G. and Hou,X.-M. (2022) Recen ad ances in applying G-quad uplex o SARS-CoV-2 a ge ing and diagnosis: a e ie w. In . J. Biol. Mac omol. , 221 , 1476–1490. 13. Puig Lomba di,E. and Londo ˜ no-Vallejo,A. (2020) A guide o compu a ional me hods o G-quad uplex p edic ion. Nucleic Acids Res. , 48 , 1–15. 14. Ruggie o,E. and Rich e ,S.N. (2020) Vi al G-quad uplexes: new on ie s in i us pa hogenesis and an i i al he apy. Annu. Rep. Med. Chem. , 54 , 101–131. 15. Teng,Y., Zhu,M., Chi,Y., Li,L. and Jin,Y. (2022) Can G-quad uplex become a p omising a ge in HBV he apy? F on . Immunol. , 13 , 1091873. 16. Chak abo y,D. and Ghosh,S. (2017) The epsilon mo i o hepa i is B i us RNA exhibi s a po assium-dependen ibonucleoly ic ac i i y. FEBS J. , 284 , 1184–1203. 17. Biswas,B., Kandpal,M. and Vi ekanandan,P. (2017) A G-quad uplex mo i in an en elope gene p omo e egula es ansc ip ion and i ion sec e ion in HBV geno ype B. Nucleic Acids Res. , 45 , 11268–11280. 18. Sa ana han,N. and Vi ekanandan,P. (2019) G-Quad uplexes: mo e Than Jus a Kink in Mic obial Genomes. T ends Mic obiol. , 27 , 148–163. 19. Meie -S ephenson,V., Badmalia,M.D., M ozowich,T., Lau,K.C.K., Schul z,S.K., Gemmill,D.L., Osiowy,C., an Ma le,G., Co in,C.S. and Pa el,T.R. (2021) Iden i ica ion and cha ac e iza ion o a G-quad uplex s uc u e in he p e-co e p omo e egion o hepa i is B i us co alen ly closed ci cula DNA. J. Biol. Chem. , 296 , 100589. 20. Fleming,A.M., Nguyen,N.L.B. and Bu ows,C.J. (2019) Colocaliza ion o m6A and G-quad uplex- o ming sequences in i al RNA (HIV, zika, hepa i is B, and SV40) sugges s opological con ol o adenosine N6-me hyla ion. ACS Cen . Sci. , 5 , 218–228. 21. Somku i,J., Moln ´ a ,O.R., G ´ ad,A. and Smelle ,L. (2021) P essu e pe u ba ion s udies o noncanonical i al nucleic acid s uc u es. Biology , 10 , 1173. 22. Moln ´ a ,O.R., V ´ egh,A., Somku i,J. and Smelle ,L. (2021) Cha ac e iza ion o a G-quad uplex om hepa i is B i us and i s s abiliza ion by binding TMPyP4, BRACO19 and PhenDC3. Sci. Rep. , 11 , 23243. 23. Sun,J., Wu,G., Pas o ,F., Rahman,N., Wang,W.-H., Zhang,Z., Me le,P., Hui,L., Sal e i,A., Du an el,D. e al. (2022) RNA helicase DDX5 enables STAT1 mRNA ansla ion and in e e on signalling in hepa i is B i us eplica ing hepa ocy es. Gu , 71 , 991–1005. 24. Nu k,S ., Ko en,S ., Rhie,A., Rau iainen,M., Bzikadze,A.V., Mikheenko,A., Vollge ,M.R., Al emose,N., U alsky,L., Ge shman,A. e al. (2022) The comple e sequence o a human genome. Science , 376 , 44–53. 25. Ok onechnik o ,K., Goloso a,O., Fu so ,M. and Team,U. (2012) Unip o UGENE: a uni ied bioin o ma ics oolki . Bioin o ma ics , 28 , 1166–1167. 26. C ooks,G .E., Hon,G ., Chandonia,J.-M. and B enne ,S.E. (2004) WebLogo: a sequence logo gene a o . Genome Res. , 14 , 1188–1190. 27. Ha kins,K.M. and S one,A.C. (2015) Ancien pa hogen genomics: insigh s in o iming and adap a ion. J. Hum. E ol. , 79 , 137–149. 28. de-Dios,T., Scheib,C.L. and Houldc o ,C.J. (2023) An adagio o i uses, played ou on ancien DNA. Genome Biol. E ol. , 15 , e ad047. 29. Taubenbe ge ,J.K., Bal imo e,D., Dohe y,P.C., Ma kel,H., Mo ens,D.M., Webs e ,R.G. and Wilson,I.A. (2012) Recons uc ion o he 1918 in luenza i us: unexpec ed ewa ds om he pas . Mbio , 3 , e00201-12. 30. Sudhan,S .S . and Sha ma,P. (2020) Human i uses: eme gence and e olu ion. Eme g. Reeme g. Vi al Pa hog. , 2020 , 53–68. 31. I ing-Pease,E.K., Muk upa ela,R., Dannemann,M. and Racimo,F. (2021) Quan i a i e human paleo gene ics: w ha can ancien DN A ell us abou complex ai e olu ion? F on . Gene . , 12 , 703541. 32. Sch ¨ adle ,S. and Hild ,E. (2009) HBV li e cycle: en y and mo phogenesis. Vi uses , 1 , 185–209. 33. Boh ´ alo ´ a,N., Can a a,A., Ba as,M., Kau a,P., ˇ S ˇ as n´y,J., Pe ˇ cinka,P., Foj a,M. and B ´ azda,V. (2021) T acing dsDNA i us-hos coe olu ion h ough co ela ion o hei G-quad uplex- o ming sequences. In . J. Mol. Sci. , 22 , 3433. 34. Ba as,M., B ´ azda,V., Boh ´ alo ´ a,N., Can a a,A., Voln ´ a,A., S achu o ´ a,T., Malacho ´ a,K., J agelsk ´ a,E.B., P o ubiako ´ a,O., ˇ Ce e ˇ n,J. e al. (2020) In-dep h bioin o ma ic analyses o nido i ales including human SARS-CoV-2, SARS-CoV, MERS-CoV i uses sugges impo an oles o non-canonical nucleic acid s uc u es in hei li ecycles. F on . Mic obiol. , 11 , 1583. 35. La igne,M., Helynck,O., Rigole ,P., Boud ia-Souilah,R., No wako wski,M., Ba on,B., B ¨ul ´ e,S., Hoos,S., Raynal,B., Gui a ,L. e al. (2021) SARS-CoV-2 Nsp3 unique domain SUD in e ac s wi h guanine quad uplexes and G4-ligands inhibi his in e ac ion. Nucleic Acids Res. , 49 , 7695–7712. 36. Ely,B. (2021) Genomic GC con en d i s downwa d in mos bac e ial genomes. PLoS One , 16 , e0244163. 37. ˇ Sma da,P ., Bu e ˇ s,P ., Ho o ´ a,L., Lei ch,I.J., Mucina,L., Pacini,E., Tich´y,L., G ulich,V. and Ro eklo ´ a,O. (2014) Ecological and e olu iona y signi icance o genomic GC con en di e si y in monoco s. P oc. Na l. Acad. Sci. U.S.A. , 111 , E4096–E4102. 38. Wang,Y., Mao,J.-M., Wang,G.-D., Luo,Z.-P., Yang,L., Yao,Q. and Chen,K.-P. (2020) Human SARS-CoV-2 has e ol ed o educe CG dinucleo ide in i s open eading ames. Sci. Rep. , 10 , 12331. 39. Ma y ´ a ˇ sek,R. and Ko a ˇ ´ ık,A. (2020) Mu a ion pa e ns o human SARS-CoV-2 and ba RaTG13 co ona i us genomes a e s ongly biased owa ds C > U ansi ions, indica ing apid e olu ion in hei Hos s. Genes (Basel) , 11 , 761. 40. Goswami,P., Ba as,M., Lexa,M., Boh ´ alo ´ a,N., Voln ´ a,A., ˇ Ce e ˇ n,J., ˇ Ce e ˇ no ´ a,V., Pe ˇ cinka,P., ˇ Spunda,V., Foj a,M. e al. (2021) SARS-CoV-2 ho -spo mu a ions a e signi ican ly en iched wi hin in e ed epea s and CpG island loci. B ie Bioin o m , 22 , 1338–1345. 41. B ´ azda,V., Po ubiako ´ a,O., Can a a,A., Boh ´ alo ´ a,N., Cou al,J., Ba as,M., Foj a,M. and Me gny,J.-L. (2021) G-quad uplexes in H1N1 in luenza genomes. BMC Genomics [Elec onic Resou ce] , 22 , 77. 42. Boh ´ alo ´ a,N., Can a a,A., Ba as,M., Kau a,P., ˇ S ˇ as n´y,J., Pe ˇ cinka,P., Foj a,M., Me gny,J.-L. and B ´ azda,V. (2021) Analyses o i al genomes o G-quad uplex o ming sequences e eal hei co ela ion wi h he ype o in ec ion. Biochimie , 186 , 13–27. 43. B own,J.C. (2007) High G+C con en o he pes simplex i us DNA: p oposed ole in p o ec ion agains e o ansposon inse ion. Open Biochem. J. , 1 , 33–42. 44. Vinog ado ,A.E. and Ana skaya,O.V. (2017) DNA helix: he impo ance o being AT- ich. Mamm. Genome , 28 , 455–464. 45. Cal ignac-Spence ,S., D¨ux,A., Goga en,J.F. and Pa ono,L.V. (2021) Chap e Two - Molecula a cheology o human i uses. In: Kielian,M., Me enlei e ,T.C. and Roossinck,M.J. (eds.) Ad ances in Vi us Resea ch . Academic P ess, Vol. 111 , pp. 31–61. C The Au ho (s) 2023. Published by Ox o d Uni e si y P ess on behal o Nucleic Acids Resea ch. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p: // c ea i ecommons.o g / licenses / by / 4.0 / ), which pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed. Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024