scieee Science in your language
[en] (orig)

G-quadruplexes in the evolution of hepatitis B virus

Abstract

Hepatitis B virus (HBV) is one of the most dangerous human pathogenic viruses found in all corners of the world. Recent sequencing of ancient HBV viruses revealed that these viruses have accompanied humanity for several millenia. As G-quadruplexes are considered to be potential therapeutic targets in virology, we examined G-quadruplex-forming sequences (PQS) in modern and ancient HBV genomes. Our analyses showed the presence of PQS in all 232 tested HBV genomes, with a total number of 1258 motifs and an average frequency of 1.69 PQS per kbp. Notably, the PQS with the highest G4Hunter score in the reference genome is the most highly conserved. Interestingly, the density of PQS motifs is lower in ancient HBV genomes than in their modern counterparts (1.5 and 1.9/kb, respectively). This modern frequency of 1.90 is very close to the PQS frequency of the human genome (1.93) using identical parameters. This indicates that the PQS content in HBV increased over time to become closer to the PQS frequency in the human genome. No statistically significant differences were found between PQS densities in HBV lineages found in different continents. These results, which constitute the first paleogenomics analysis of G4 propensity, are in agreement with our hypothesis that, for viruses causing chronic infections, their PQS frequencies tend to converge evolutionarily with those of their hosts, as a kind of 'genetic camouflage' to both hijack host cell transcriptional regulatory systems and to avoid recognition as foreign material.

Read accessible full text

G-quadruplexes in the evolution of hepatitis B virus

Author: Brázda, Václav; Dobrovolná, Michaela; Bohálová, Natália; Mergny, Jean-Louis
Publisher: Oxford University Press
Year: 2023
DOI: 10.1093/nar/gkad556
Source: https://dspace.vut.cz/bitstreams/d5c5eba7-36b6-4b67-9432-fab5b22152e5/download
7198–7204 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 Published online 3 July 2023
h ps://doi.o g/10.1093/na /gkad556
G-quad uplexes in he e olu ion o hepa i is B i us
V
´
acla B
´
azda
1 ,
*
, Michaela Dob o oln
´
a
1 , 2
, Na
´
alia Boh
´
alo
´
a
1 and Jean-Louis Me gny
1 , 3 ,
*
1
Ins i u e o Biophysics o he Czech Academy o Sciences, B no, Czech Republic,
2
Facul y o Chemis y, B no
Uni e si y o Technology, Pu ky
ˇ
no a 118, 612 00 B no, Czech Republic and
3
Labo a oi e d’Op ique e Biosciences
(LOB), Ecole Poly echnique, CNRS, INSERM, Ins i u Poly echnique de Pa is, 91120 Palaiseau, F ance
Recei ed Ma ch 30, 2023; Re ised May 23, 2023; Edi o ial Decision June 14, 2023; Accep ed June 19, 2023
ABSTRACT
Hepa i is B i us (HBV) is one o he mos dange ous
human pa hogenic i uses ound in all co ne s o he
wo ld. Recen sequencing o ancien HBV i uses e-
ealed ha hese i uses ha e accompanied human-
i y o se e al millenia. As G-quad uplexes a e con-
side ed o be po en ial he apeu ic a ge s in i ol-
ogy, we examined G-quad uplex- o ming sequences
(PQS) in mode n and ancien HBV genomes. Ou
anal yses sho wed he p esence o PQS in all 232
es ed HBV genomes, wi h a o al numbe o 1258 mo-
i s and an a e a ge equenc y o 1.69 PQS pe kbp.
No ably, he PQS wi h he highes G4Hun e sco e in
he e e ence genome is he mos highly conse ed.
In e es ingly, he densi y o PQS mo i s is lowe in
ancien HBV genomes han in hei mode n coun e -
pa s (1.5 and 1.9 / kb, espec i ely). This mode n e-
quency o 1.90 is e y close o he PQS equency o
he human genome (1.93) using iden ical pa ame e s.
This indica es ha he PQS con en in HBV inc eased
o e ime o become close o he PQS equency in
he human genome. No s a is ically signi ican di -
e ences we e ound be ween PQS densi ies in HBV
lineages ound in di e en con inen s. These esul s,
which cons i u e he i s paleogenomics analysis o
G4 p opensi y, a e in ag eemen wi h ou hypo he-
sis ha , o i uses causing ch onic in ec ions, hei
PQS equencies end o con e ge e olu iona il y wi h
hose o hei hos s, as a kind o ‘gene ic camou lage’
o bo h hijack hos cell ansc ip ional egula o y sys-
ems and o a oid ecogni ion as o eign ma e ial.
GRAPHICAL ABSTRACT
INTRODUCTION
Hepa i is B i us (HBV) belongs o he genus O hohep-
adna i us; i s genome is cons i u ed o double-s anded
DNA. This i us causes Hepa i is B , a highly con a-
gious, po en ially a al disease ha a ec s an es ima ed
257 million people wo ldwide, esul ing in an es ima ed
820 000 dea hs e e y yea ( h ps://www.who.in /news- oom/
ac - shee s/de ail/hepa i is- b ). This i us is pa icula ly se-
e e, as a pp oxima el y one in i e ca ie s die om ci ho-
sis and / o de elop hepa ocellula ca cinoma. HBV is ans-
mi ed p ima ily h ough blood and body luids and he
incuba ion pe iod is a iable, usually be ween 30 and 180
days. Du ing eplica ion, HBV DNA o ms a minich omo-
some in he nucleus o in ec ed hepa ocy es ( 1 , 2 ) and i s
genome is eplica ed h ough a p ocess o e e se ansc ip-
ion o he key in e media e p e-genomic RNA in hepa o-
cy es, which is also an mRNA empla e o he HBV p o-
eins ( 3 , 4 ). Hepadna i uses in ec ing o he hos s ha e e-
cen ly been iden i ied, including ba s , ogs , liza ds , ish, and
he capuchin monkey ( 5–8 ). Analyses o ancien genomes
ha e e ealed ha he mos ecen common ances o o all
HBV lineages is es ima ed o ha e exis ed be ween ∼20 000
*
To whom co espondence should be add essed. Tel: +420 541517231; Email: acla @ibp.cz
Co espondence may also be add essed o Jean-Louis Me gny. Tel: +33 766290967; Email: jean-louis.me gny@poly echnique.edu
C
The Au ho (s) 2023. Published by Ox o d Uni e si y P ess on behal o Nucleic Acids Resea ch.
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p: // c ea i ecommons.o g / licenses / by / 4.0 / ), which
pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed.
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7199
and 12 000 yea s ago, and he i us was ound o be p esen
in Eu opean and Sou h Ame ican hun e -ga he e s du ing
he ea ly Holocene pe iod ( 9 ).
A la ge amoun o li e a u e is de o ed o he occu ence
o G-quad uplexes in i uses, especially o he possibili y o
using hese s uc u es in he apy. Comp ehensi e bioin o -
ma ics analyses ha e aced pu a i e G4- o ming sequences
in he genome o almos all human i uses, showing ha
hei dis ibu ion and p esence a e highly conse ed. The e-
o e, hese DNA o RNA s uc u es can be a sui able a -
ge o a ge ed he apy. Some G-quad uplex ligands ha e
been shown o ha e an i i al ac i i y, o example, agains
HIV ( 10 ), he pes simplex i us I (HSV-1) ( 11 ), SARS-CoV-
2 ( 12 ) and o he s. A comp ehensi e analysis o all sequenced
i uses ha ha e a la en phase in hei li e cycle showed
ha hei G-quad uplex con en is co ela ed wi h ha o
he hos ( 13 ). In con as , i uses causing acu e in ec ions
wi hou a la en phase end o elimina e G-quad uplex se-
quences, as hey can become oadblocks du ing eplica ion,
ansc ip ion and / o e e se ansc ip ion ( 14 ).
Mo e speci ically, G4s ha e been ound o be ele an
in HBV in ec ion, bo h a he DN A and RN A le el ( 15 ).
Chak abo y and Ghosh epo ed ha an RNA sequence
p esen in HBV RNA exhibi ed a sequence-independen
ans-ac ing nuclease ac i i y, and ha his sequence adop s
a G4 con o ma ion ( 16 ). Biswas e al. analyzed a G4 p one
mo i ( GGGAGTGGGAGCATTCGGGCCAGGG ) ha is highly
conse ed only in HBV geno ype B, and was shown o
adop a hyb id s uc u e ( 17 ). In e es ingl y, m u a ions dis-
up ing his G-quad uplex in HBV geno ype B cons uc s
we e associa ed wi h impai ed i ion sec e ion. The au ho s
p oposed ha his G4 media es enhancemen o ansc ip-
ion and i ion sec e ion in his HBV geno ype. In a la e e-
iew, hey no e ha among i uses con aining a G4 in hei
genome, hose associa ed wi h cance a e o e - ep esen ed,
including HBV ( 18 ).
In con as , some G4s end o be conse ed in all geno-
ypes, as epo ed by Meie -S ephenson e al ( 19 ) o a
DNA sequence ound in he p e-co e p omo e egion
( CTGGGAGGAGCTGGGGGAGGAGA ). They demons a ed a
ole o his quad uplex in i al eplica ion by compa ing
he wild- ype mo i o non-G4- o ming mu an s in i o .
Fleming e al iden i ied conse ed po en ial G4 sequences
in se e al i al genomes ele an o human heal h, and
showed ha hese mo i s can p o ide a ame wo k o N6-
me hyladenosine (m6A) ins alla ion wi hin he loops o
RNA G4 sequences ound in se e al i uses, including HBV
( 20 ). Somku i e al. de e mined olume changes in h ee
HBV G4 s uc u es using biophysical app oaches in i o .
They in es iga ed h ee DNA sequences: GGCTGGGGCTTG-
GTCATGGGCCATCAG , GGGAGTGGGAGCATTCGGGCCAGGG
and TTGGGTGGCTTTGGGGCATGGAC ( 21 ). The same g oup
in es iga ed one o hese sequences in mo e de ails (HepB;
GGCTGGGGCTTGGTCATGGGCCATCAG , ound in he coding
egion o he polyme ase p o ein), and analyzed i s in e -
ac ion wi h G4 ligands in i o ( 22 ). Finally, Sun e al. e-
cen ly in es iga ed he ole o cellula G4s (mo i s ound in
he hos cell genome) in HBV in ec ion, demons a ing ha
he DDX5 helicase, known o be capable o esol ing RNA
G4 s uc u es, is a key egula o o he in e e on (IFN)
esponse agains his i us. DDX5 down egula ion is ob-
se ed du ing HBV eplica ion and in poo p ognosis HBV-
ela ed hepa ocellula ca cinoma (HCC) ( 23 ). All o hese
esul s poin ou he links be ween i al (o hos ) DNA and
RNA G4s and HBV.
In his pape , we ha e analyzed 232 HBV genomes om
samples co e ing a mo e han 10- housand-yea his o y o
he p esence o G-quad uplex o ming sequences. Ou e-
sul s show an e olu iona y shi o an inc eased numbe o
G-quad uplex es in ecen HBV i uses, poin ing o he im-
po ance o G-quad uplexes in he HBV li e-cycle in human
li e cells.
MATERIALS AND METHODS
Genomes
232 HBV alignmen s we e downloaded om he supple-
men a y ma e ials a (9). Sequences we e ob ained o 122
mode n geno ypes and di e en g oups o 110 ancien
s ains, di ided in o g oups based on mPTP classi ica ion.
As he e e ence genome, we ook NC 003977.2 and, o
analyses o phylogene ically ela ed i uses wi h hos s o he
han human, we il e ed e e ence genomes om O hohep-
adna i uses. In o al, we downloaded 21 addi ional HBVs
ha ing a non-human hos , in ec ing bi ds (7 genomes), ba s
(5 genomes), ish (2 genomes), o he Mammals (i.e. nei he
human no ba s; 6 genomes) and amphibians (1 genome, Ti-
be an og hepa i is B i us). HBV G4 con en s we e com-
pa ed o he gapless human genome, he new elome e- o-
elome e assembly o he human genome ( 24 ), which was
downloaded om NCBI (T2T-CHM13 2.0).
G4Hun e analyses
All sequences we e analyzed using G4Hun e
( h p://bioin o ma ics .ibp .cz ) o iden i y PQS sequences.
G4Hun e ’s de aul pa ame e s we e used (25 nucleo ides
o window size and 1.2 o h eshold). These se ings ha e
p e iously been shown o iden i y expe imen ally- alida ed
quad uplex s uc u es. The lis o all o ganisms es ed
and he esul s o he analyses we e downloaded om he
supplemen a y ma e ials a ( 9 ).
S a is ical e alua ion
Da a wi h G4Hun e esul s we e me ged in an Excel ile o
s a is ical e alua ion. G4Hun e esul s , leng hs , and GC
con en o analyzed sequences a e accessible in Supplemen-
a y ma e ial 01. A sca e plo was gene a ed in G aph-
Pad P ism ( 8.0.1), Violin plo s we e cons uc ed in R (
4.2.0) wi h ggplo 2. S a is ical signi icance was es ed using
S uden ’s T- es . No mali y o da a was de e mined using
Shapi o-Wilk es .
Cons uc ion o LOGO sequence
All sequences o ancien and mode n HBV genomes we e
uploaded in o UGENE so wa e ( 25 ) and he loca ion o
PQS sequences we e ex ac ed using Clus alW alignmen .
LOGO sequence was gene a ed in aligned sequences and
WebLogo 3 ool ( 26 ).
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
7200 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14
Table 1. S a is ic da a o G4Hun e analyses o HBV i uses. n Seq (numbe o s ains), Leng h (leng h o he sequence, n ), GC % (a e age GC con en ),
PQS n ( o al numbe o p edic ed PQS wi h a G4Hun e sco e o 1.2 o mo e), Mean PQS (a e age numbe o p edic ed PQS), Min PQS (lowes equency
o p edic ed PQS), Max PQS (highes equency o p edic ed PQS), PQS pe 1000 GC (PQS equency pe 1000 GC)
G oups Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC
All HBV 232 3214 45.8 1258 1.69 0.61 4.04 3.66
G oups Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC
ancien 122 3217 43.1 587 1.50 0.61 2.83 3.46
mode n 110 3210 48.8 671 1.90 0.94 4.04 3.88
H. sapiens T2T 1* 3.05 ×10
9 41.6 5.45 10
6 1.93 - - 4.72
Subg oups** Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC
A 21 3227 47.0 105 1.55 2.00 8.00 3.30
B 14 3222 49.0 120 2.66 1.55 4.04 5.43
C 16 3222 49.1 96 1.86 1.55 3.10 3.79
D 57 3190 47.7 317 1.74 0.94 2.83 3.66
E 3 3218 48.4 25 2.59 2.17 2.80 5.35
F 13 3222 48.8 81 1.93 1.24 2.48 3.96
G 3 3255 48.2 12 1.23 1.23 1.23 2.55
H 5 3222 48.9 39 2.42 1.86 3.11 4.94
I 3 3219 48.8 20 2.07 1.55 2.49 4.24
J 1 3187 49.0 4 1.26 1.26 1.26 2.56
Ancien Ame ican 4 3214 38.4 15 1.17 0.93 1.55 3.12
Ea ly Ana olian a me 1 3189 33.7 2 0.63 0.63 0.63 1.86
Mesoli hic 8 3192 45.8 33 1.29 0.63 2.20 3.20
WENBA 65 3229 42.1 288 1.37 0.61 2.20 3.27
o u 2 3191 49.1 8 1.25 1.25 1.25 2.56
gbn 3 3194 49.1 20 2.09 1.57 2.51 4.25
czp 3 3191 48.3 18 1.88 1.25 2.51 3.89
o he 10 3211 45.2 55 1.71 0.62 4.04 3.81
*Comple e elome e- o- elome e human genome (22 + X + Y ch omosomes)
**The g oups we e de ined by Koche e al. ( 10 ) and is based on he mul i- a e Poisson T ee P ocesses (mPTP) as a esul s o gene ic clus e s numbe
conside ing a phylogene ic inpu ee ( 26 ).
RESULTS
We analyzed he p esence o PQS in 232 HBV genomes
(122 ancien and 110 ecen ) using G4Hun e . All human
HBV genomes a e simila in leng h, a ying om 3180 o
3300 bp. Compa isons o ancien and ecen samples show
a sligh and non-signi ican change in a e age leng h om
3217 o 3210 bp. On he o he hand, hese genomes a y in
GC con en , and we ound he p esence o G-quad uplex-
o ming sequences in all HBV genomes in he da ase . In
o al, we ound 1258 PQS wi h a mean equency o 1.69
PQS pe kbp (Table 1 ). The mean equency o PQS in
ancien genomes is 1.50 / kb, compa ed wi h he mean e-
quency in mode n HBV genomes o 1.90. Fo compa ison
we also analyzed he ne wly pub lished human gapless assem-
bly. The PQS equency in he human genome is 1.93, which
is almos iden ical o he a e age PQS equency o mode n
HBV genomes (Table 1 ). Al hough he e is no signi ican
change in he leng h o ancien and mode n HBV genomes,
compa ison o PQS densi y shows ha he mode n i uses
a e subs an ially iche in PQS (Figu e 1 ) ( P - alue = 3.6e-
08). The mode n HBV genomes no only ha e a highe PQS
equency, bu also ha e a highe GC con en . To e alua e i
he change in PQS equency is s a is ically signi ican a e
aking in o accoun GC con en , we ecalcula ed he PQS
equency acco ding o GC con en (Table 1 , las column).
E en a e his co ec ion, he PQS equency / GC con en
is highe in mode n han in ancien HBV genomes (3.88 e -
sus 3.46 pe housand GC o he mode n and ancien HBV
genomes, espec i ely: P - alue = 1.8e-03).
We hen u he di ided genomes acco ding o di e -
en geno ypes. While mos cu en geno ypes ha e an a -
Figu e 1. Compa ison o PQS equencies in ancien and mode n HBV
genomes.
e age PQS equency highe han 2 PQS / kb, ou o he
i e ancien genomes ha e PQS equencies lowe han his
alue. The highes PQS equencies we e ound in mod-
e n geno ypes B, E and H, he lowes in ancien Ame ican,
Mesoli hic and o he ancien geno ypes (Figu e 2 ).
We p esen he PQS equency pe kb ( o all mo i s wi h
a G4Hun e sco e > 1.2) and he leng h o he genome as a
unc ion o ime o each ancien sequence (Supplemen a y
ma e ial 02). As men ioned be o e, he longes sequence
only di e s om he sho es by 120 bp, and he leng h o
he genome does no change signi ican ly o e ime (Pea -
son = −0.1387, P ( wo- ailed) = 1.6e-01). Unlike genome
leng h, PQS equencies pe kbp we e ound o inc ease o e
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7201
Figu e 2. PQS equency o HBV subg oups di ided in o ancien and mode n geno ypes. Ancien Ame ican 1, 2, 3 ca ego ies we e me ged. Mesoli hic 1
and 2 geno ypes we e me ged.
Table 2. Posi ion o G4 sequences in he human HBV e e ence genome (NC 003977.2) and i s loca ion conse a ion. A nega i e sco e co esponds o a
C- ich sequence, meaning ha i is i s complemen a y s and which will be G4-p one
Posi ion S and Sequence G4H sco e Fea u e / loca ion G4 conse a ion
316 −CCCC AA CC T CC AAT C A C T C A CC −1.41 S, DNA polyme ase
N- e minal domain 84.5%
737 + GG AT G AT G T GG TATT GGGGG +1.50 S, DNA polyme ase
C- e minal domain 70.3%
1133 −CC TGAA CC TTTA CCCC GTTG CCC −1.30 P, DNA polyme ase
C- e minal domain 19.4%
1722 + GGG A GG AGTT GGGGG A GG A G ATTA GG +1.65 T ansac i a ion p o ein X 92.2%
1887 + GGG T GG CTTT GGGG CAT GG +1.63 C , Hepa i is co e p o ein 40.5%
ime (Pea son = 0.2903, P ( wo- ailed) = 2.8e-03; Supple-
men a y ma e ial 02).
The whole genome o HBV is ansc ibed in o one long
p e-genomic RN A (mRN A-pgRN A) ha encodes all HBV
p o eins, and pgRNA also ac s as a empla e o e e se
ansc ip ion. Analysis o PQS localiza ion in he HBV e -
e ence genome (NC 003977.2) shows ha all PQS a e lo-
ca ed in he icini y o gene egions, which is no su p ising
conside ing ha he HBV genome is small and he en i e
genome is used e y e ec i ely o p oduce he ew p o eins
necessa y o i s unc ion, such as DN A pol yme ase, ans-
ac i a ion and capsid p o eins (Table 2 ).
The PQS wi h he highes G4Hun e sco e in he HBV
e e ence genome is loca ed su ounding posi ion 1722, in
he egion ha codes o ansac i a ion p o ein X. In e -
es ingly, a PQS is p esen in almos all HBV genomes a
his loca ion (214 o 232; o 92.2% o HBV genomes an-
alyzed), making i he mos conse ed mo i be ween all
PQS in he e e ence HBV genome. Conse a ion implies
ha his sequence posi ion has been main ained by selec-
i e p essu e. Compa ison o he LOGO sequence o his
loca ion in mode n and ancien HBV genomes (Figu e 3 )
demons a es ha a G- ich mo i is p ese ed in all s ains.
Ne e heless, his ‘G- ichness’ is e en mo e s iking in mod-
e n compa ed o ancien s ains, wi h Gs becoming p edom-
inan a posi ions 8 and 11 (Figu e 3 , a ows), while ancien
Figu e 3. L OGO ep esen a ion o he consensus mo i ound a ound po-
si ion 1722 (acco ding o he e e ence HBV genome) in mode n and an-
cien HBV genomes. Guanine nucleo ides p edomina e in mode n s ains
a posi ions 8 and 11 (a ows) while in ancien HBV genomes o he nu-
cleo ides a e p esen (T and A in posi ion 8). A less s iking G-en ichmen
is also ound a posi ions 14 and 15. As a consequence, while bo h con-
sensus mo i s a e compa ible wi h G4 o ma ion, he mode n sequence is
mo e a o able han he ancien one (G4hun e sco es o 2.11 and 1.83,
espec i ely).
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
7202 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14
Table 3. G4Hun e sco e in HBV i uses, g ouped by con inen and age
G oups Seq Leng h Mean PQS Min PQS Max PQS
All 232 3214 1.69 0.61 4.04
Ancien Seq Leng h Mean PQS Min PQS Max PQS
Aus alia 0 - - - -
Ame ica 7 3217 1.33 0.93 3.11
A ica 1 3229 0.62 - -
Asia 34 3219 1.54 0.61 2.51
Eu ope 82 3217 1.50 0.61 2.83
Mode n Seq Leng h Mean PQS Min PQS Max PQS
Aus alia 14 3215 1.93 0.94 3.10
Ame ica 31 3216 1.96 0.94 3.11
A ica 15 3205 1.83 1.24 2.80
Asia 110 3206 1.93 0.94 4.04
Eu ope 12 3209 1.66 1.23 2.48
HBV genomes mo e o en exhibi T / A o A a hese posi-
ions. A highe –– nea 100% –– p e alence o Gs is also is-
ible a posi ions 14 and 15 (Figu e 3 ) in mode n genomes.
O e all, while bo h ancien and mode n consensus mo i s
a e G4-p one, he mode n sequence is mo e a o able, as
shown by he highe G4Hun e sco e.
We also di ided HBV genomes acco ding o he geo-
g aphic place o sampling. Mos o he ancien HBV sam-
ples we e ound in Eu ope, only one in A ica and none in
Aus alia. None heless, ancien genomes ha e a lowe PQS
equency han mode n HBV genomes, ega dless o hei
con inen o o igin (Table 3 ; also see G aphical Abs ac ).
DISCUSSION
Pa hogens e ol e in esponse o human biological changes
alongside sociocul u al and echnological de elopmen s
( 27 ). Ancien i al genomes p o ide in o ma ion on he
e olu ion o i uses o e bo h ime and space and p o-
ide insigh in o he changes ha may ha e occu ed in
i ulence and ansmissibili y ( 28 ). Cu en ad anced ech-
niques o isola ion o nucleic acids and sequencing ha e al-
lowed paleogenomic o ‘a cheo i ology’ in es iga ions. The
in amous 1918 ‘Spanish’ in luenza pandemic was he sou ce
o he i s ancien pa hogen genome ( 29 ) and a cheo i ol-
ogy has been g owing apidly since hen. Conside ing ha
wo- hi ds o all human pa hogens a e i uses ( 30 ), paleo-
gene ics o i al genomes p o ides an in e es ing iewpoin
on human his o y ( 31 ).
While se e al ancien i al genomes a e a ailable, he
mos comp ehensi e da ase deals wi h ancien HBV
genomes ( 28 ). Wi h mo e han 200 million people su e ing
om ch onic HBV in ec ion, HBV can be conside ed as a
common i us. HBV has a li e-cycle ha equi es i s double-
s anded DNA genome o each he hos cell nucleus ( 32 ),
in con as wi h RNA i uses such as in luenza o SARS-
CoV-2 ha cause only acu e in ec ions and con ain RNA
genomes ha can be eplica ed and ansla ed in he cy o-
plasm. Vi uses causing acu e in ec ions end o ha e a low
PQS equency, while G4s end o be ubiqui ous in mos o -
ganisms.
G4s in he HBV genome a e impo an s uc u es ha
egula e ansc ip ion and i ion sec e ion in HBV geno-
ype B ( 15 , 17 ). In his epo , we analyzed he p esence o
PQS in mul iple HBV genomes, om ancien o cu en
HBV s ains. Impo an ly, we ound ha PQS equency is
highe in ecen compa ed o ancien s ains. I was shown
p e iously ha he G4 equency o dsDNA i uses co e-
la es wi h he PQS equency o he hos , shown o dsDNA
i uses in ec ing A chaea, Bac e ia and Euka yo a ( 33 ). In
ag eemen wi h hose da a, we ound ha he densi y o G4
mo i s in mode n HBV s ains (dsDNA i uses ha also
expe ience a la en phase) ends o con e ge o he o e all
G4 densi y o he human genome. We p opose ha mu a-
ions which led o ‘PQS in eg a ion’ (new G4 mo i s wi hin
he HBV genome) we e e olu iona y p e e ed and ixed in
HBV pa hogenic s ains du ing e olu ion. I seems ha he
opposi e p ocess may occu in i uses causing acu e in ec-
ions, as ound o SARS-CoV-2, whe e he PQS equency
is ex emely low compa ed o he PQS equencies o o he
co ona i uses ( 34 , 35 ).
Vi uses ha e highly a iable genomes and a e p one o
mu a ions, in con as o cellula and especially mul icellu-
la o ganisms, as desc ibed epea edly. This is especially ue
o RN A i uses, w he e he m u a ion a e is se e al o de s
o magni ude highe han in DNA-based genomes. A d a-
ma ic dec ease in GC con en has been desc ibed o se e al
bac e ial species ( 36 ) and o some plan species wi h holo-
cen ic ch omosomes ( 37 ). Simila ly, a apid dec ease in GC
con en o e ime was ound o some i uses causing acu e
in ec ions (and wi hou la ency connec ed o nuclea local-
iza ion), including Nido i ales, in luenza genomes and he
con empo a y SARS-CoV-2 ou b eak ( 38 ). Compa ison o
SARS-CoV-2 genomes showed a s ong p e e ence o mu-
a ions in GC islands and C > U ansi ions, leading o a
dec ease in G4 p opensi y ( 39 , 40 ). As a consequence, he
PQS equency o hese i uses causing only acu e in ec-
ions is gene ally e y low (0.03 o SARS-2 ( 35 ), 0.56 o
in luenza H1N1 genomes ( 41 )).
The opposi e end is ue o HBV, which exhibi s a la-
en s a e and main ains i s genome in he nucleus: mode n
HBV genomes ha e a signi ican ly highe PQS equency
compa ed o ancien HBV genomes. Ou esul s a e in line
wi h a b oad s udy compa ing PQS equencies in i uses
wi h p edominan ly pe sis en o acu e ypes o in ec ion
( 42 ). Acco ding o ha s udy, i uses causing pe sis en in-
ec ion a e en iched in PQS compa ed o acu ely in ec ious
i uses. Impo an ly, his obse a ion is also alid wi hin
i uses causing hepa i is: HAV (hepa i is A i us - caus-
ing acu e in ec ion) ha e a low PQS equency, while he
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024

Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7203
Hepadna i idae ha cause ch onic in ec ions ( o which hu-
man HBV belongs) ha e a signi ican ly highe PQS e-
quency ( 42 ).
The p esence o G4s depends on guanine con en in he
genome, and one o he possible ad an ages o a GC- ich
genome is o p o ide addi ional gene egula ion oppo u-
ni ies. In his espec , G4s ha e been shown o be impo -
an o ansc ip ion in highe o ganisms. An inc ease in
GC con en has also been documen ed in plan species ha
can g ow in seasonally cold clima es, possibly indica ing
an ad an age o GC- ich DNA du ing cell eezing, and
he genomic adap a ions associa ed wi h changing GC con-
en a e sugges ed o g ass-domina ed biomes du ing he
Te ia y pe iod ( 37 ). G4s a e o en o e ep esen ed in he
p omo e egions o highe euka yo es and ha e also been
demons a ed o con ibu e o di ec ed genome edi ing in
nema odes. Fo i uses expe iencing a la en phase in he
nucleus, ha ing a simila genome o ganiza ion as he hos
is ad an ageous o bo h a oid ecogni ion as unusual ( o -
eign) DNA, and o hijack he hos egula o y machine y. In
addi ion, he high GC con en o HSV DNA is sugges ed o
ac as a p o ec i e ea u e agains e o ansposon inse ion
( 43 ).
As AT- ich egions in humans a e mos ly associa ed wi h
condensed ch oma in ( 44 ), he shi o GC- ich i uses
could be impo an o i uses wi h la en phase o ha e a
be e chance o being ac i e in he u u e. HBV has a la-
en pe iod, he e o e, he e olu iona y p essu e o inc ease
GC con en and PQS p esence could be e olu iona y a-
o ed. In his model, he o iginal (non-human) p e-HBV
hos could ha e had a lowe G4 equency – adap a ion o
he human hos may ha e led o a hos -pa hogen PQS con-
e gence and a concomi an inc ease in G4 densi y in he
i us.
CONCLUSION
We pe o med he i s paleogenomic analysis o G4
p opensi y, applied he e o he Hepa i is B i us. We ound
ha he densi y o PQS mo i s inc eased o e ime, as i
is highe in mode n han ancien HBV genomes. The e-
quency in mode n i uses is now e y close o ha o he
human genome. This s udy should pa e he way o he
paleo genomics anal ysis o G4 sequences (and o he sim-
ila mo i s) in o he pa hogens. Un o una ely, his o ical
in o ma ion abou i uses genomes is only a ely a ail-
able: A cheo i ology is a nascen ield ( 45 ), which aces
he same obs acles as mode n genomics, bu wi h he ad-
di ional p oblem o analyzing pa ially deg aded DNA. As
no ed by he au ho s ( 45 ), HBV is ‘an excellen a ge o
he eco e y o ancien sequences due o i s ela i ely s ab le,
pa ially double-s anded ci cula DNA genome, i s high
p e alence in he human popula ion, and p olonged high
i emia du ing ch onic in ec ion’. Double-s anded DNA
is gene ally be e p ese ed han single-s anded DNA o
RNA. This may explain why HBV is cu en ly he only in-
s ance in which sequence da a a e a ailable o o e 100 an-
cien i uses. In o he si ua ions, such as a iola i us, only
a ew ancien genomes a e a ailable ( he e a e only 4, 10
and 11 HSV-1, Pa o i uses and VA RV ancien sequences
a ailab le, espec i el y ( 28 ). Anal yses o HIV-1 o he 1918
in luenza i uses is possible, bu only o e a limi ed pe iod
o ime.
A cheo i ology will be use ul o iden i y polymo phisms
impo an o human adap a ion o pa hogens, and ice-
e sa du ing he complex i us-hos ela ionships c ucial o
con inued i al p e alence ( 28 ). Mo e paleogenomic da a
will be needed o es he hypo hesis ha , o i uses causing
ch onic in ec ions, hei PQS equencies end o con e ge
e olu iona ily wi h hose o hei hos . In pa icula , gi en
ecen se ious ou b eaks, we hope ha he analysis o an-
cien i al pa hogens will p o ide c i ical knowledge abou
he na u e o new i al diseases.
DA T A A V AILABILITY
All da a a e a ailable in he manusc ip and supplemen a y
iles. HBV sequences we e uploaded om he supplemen-
a y ma e ials a ( 9 ).
SUPPLEMENT ARY DA T A
Supplemen a y Da a a e a ailable a NAR Online.
ACKNOWLEDGEMENTS
The au ho s hank P. Coa es o English p oo eading and
aluable commen s, and L. Gui a (LOB) o help ul dis-
cussions.
FUNDING
Agence de l’Inno a ion de D
´
e ense (AID) ia he Cen e
In e disciplinai e d’E udes pou la D
´
e ense e la S
´
ecu i
´
e
(CIEDS) [p ojec 2023 - Pa hogens]; ANR G4Access
[ANR-20-CE12-0023]; INCa G4Access g an s ( o J.L.M.);
SYMBIT p ojec [CZ.02.1.01 / 0.0 / 0.0 / 15 003 / 0000477] i-
nanced om he ERDF; Czech Science Founda ion [22-
21903S o V.B.]. Funding o open access cha ge: Academ y
o Sciences.
Con lic o in e es s a emen . None decla ed.
REFERENCES
1. Bock,C.T., Sch anz,P., Sch
¨
ode ,C.H. and Zen g a ,H. (1994)
Hepa i is B i us genome is o ganized in o nucleosomes in he
nucleus o he in ec ed cell. Vi us Genes , 8 , 215–229.
2. Newbold,J.E., Xin,H., Tencza,M., She man,G., Dean,J., Bo w den,S.
and Loca nini,S. (1995) The co alen ly closed duplex o m o he
hepadna i us genome exis s in si u as a he e ogeneous popula ion o
i al minich omosomes. J. V i ol. , 69 , 3350–3357.
3. Summe s,J. and Mason,W.S. (1982) Replica ion o he genome o a
hepa i is B–like i us by e e se ansc ip ion o an RNA
in e media e. Cell , 29 , 403–415.
4. Tiollais,P., Pou cel,C. and Dejean,A. (1985) The hepa i is B i us.
Na u e , 317 , 489–495.
5. MacDonald,D.M., Holmes,E.C., Lewis,J.C. and Simmonds,P. (2000)
De ec ion o hepa i is B i us in ec ion in wild-bo n chimpanzees
(Pan oglody es e us): phylogene ic ela ionships wi h human and
o he p ima e geno ypes. J. Vi ol. , 74 , 4253–4257.
6. D exle ,J.F., Geipel,A., K
¨
onig,A., Co man,V.M., an Riel,D.,
Leij en,L.M., B eme ,C.M., Rasche,A., Co on ail,V.M.,
Maganga,G.D. e al. (2013) Ba s ca y pa hogenic hepadna i uses
an igenically ela ed o hepa i is B i us and capable o in ec ing
human hepa ocy es. P oc. Na l. Acad. Sci. U.S.A. , 110 , 16151–16156.
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
7204 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14
7. Laube ,C., Sei z,S., Ma ei,S., Suh,A., Beck,J., He s ein,J., B
¨
o old,J.,
Salzbu ge ,W., Kade ali,L., B iggs,J.A.G. e al. (2017) Deciphe ing
he o igin and e olu ion o hepa i is B i uses by means o a amily o
non-en eloped ish i uses. Cell Hos Mic obe , 22 , 387–399.
8. de Ca alho Dominguez Souza,B.F., K
¨
onig,A., Rasche,A., de
Oli ei a Ca nei o,I., S ephan,N., Co man,V.M., Roppe ,P.L.,
Goldmann,N., Keppe ,R., M¨ulle ,S.F. e al. (2018) A no el hepa i is
B i us species disco e ed in capuchin monkeys sheds new ligh on
he e olu ion o p ima e hepadna i uses. J. Hepa ol. , 68 , 1114–1122.
9. Koche ,A., Papac,L., Ba que a,R., Key,F.M., Spy ou,M.A.,
H¨uble ,R., Roh lach,A.B., A on,F., S ahl,R., Wissgo ,A. e al.
(2021) Ten millennia o hepa i is B i us e olu ion. Science , 374 ,
182–188.
10. Pe one,R., Bu o skaya,E., Daelemans,D., Pal`u,G., Pannecouque,C.
and Rich e ,S.N. (2014) An i-HIV-1 ac i i y o he G-quad uplex
ligand BRACO-19. J. An imic ob. Chemo he . , 69 , 3248–3258.
11. F asson,I., Sold
`
a,P., Nadai,M., Tassina i,M., Scalab in,M.,
Gokhale,V., Hu ley,L.H. and Rich e ,S.N. (2022)
Quindoline-de i a i es display po en G-quad uplex-media ed
an i i al ac i i y agains he pes simple x i us 1. An i i al R es. , 208 ,
105432.
12. Zhai,L.-Y., Su,A.-M., Liu,J.-F., Zhao,J.-J., Xi,X.-G. and Hou,X.-M.
(2022) Recen ad ances in applying G-quad uplex o SARS-CoV-2
a ge ing and diagnosis: a e ie w. In . J. Biol. Mac omol. , 221 ,
1476–1490.
13. Puig Lomba di,E. and Londo
˜
no-Vallejo,A. (2020) A guide o
compu a ional me hods o G-quad uplex p edic ion. Nucleic Acids
Res. , 48 , 1–15.
14. Ruggie o,E. and Rich e ,S.N. (2020) Vi al G-quad uplexes: new
on ie s in i us pa hogenesis and an i i al he apy. Annu. Rep. Med.
Chem. , 54 , 101–131.
15. Teng,Y., Zhu,M., Chi,Y., Li,L. and Jin,Y. (2022) Can G-quad uplex
become a p omising a ge in HBV he apy? F on . Immunol. , 13 ,
1091873.
16. Chak abo y,D. and Ghosh,S. (2017) The epsilon mo i o hepa i is B
i us RNA exhibi s a po assium-dependen ibonucleoly ic ac i i y.
FEBS J. , 284 , 1184–1203.
17. Biswas,B., Kandpal,M. and Vi ekanandan,P. (2017) A G-quad uplex
mo i in an en elope gene p omo e egula es ansc ip ion and i ion
sec e ion in HBV geno ype B. Nucleic Acids Res. , 45 , 11268–11280.
18. Sa ana han,N. and Vi ekanandan,P. (2019) G-Quad uplexes: mo e
Than Jus a Kink in Mic obial Genomes. T ends Mic obiol. , 27 ,
148–163.
19. Meie -S ephenson,V., Badmalia,M.D., M ozowich,T., Lau,K.C.K.,
Schul z,S.K., Gemmill,D.L., Osiowy,C., an Ma le,G., Co in,C.S.
and Pa el,T.R. (2021) Iden i ica ion and cha ac e iza ion o a
G-quad uplex s uc u e in he p e-co e p omo e egion o hepa i is B
i us co alen ly closed ci cula DNA. J. Biol. Chem. , 296 , 100589.
20. Fleming,A.M., Nguyen,N.L.B. and Bu ows,C.J. (2019)
Colocaliza ion o m6A and G-quad uplex- o ming sequences in i al
RNA (HIV, zika, hepa i is B, and SV40) sugges s opological con ol
o adenosine N6-me hyla ion. ACS Cen . Sci. , 5 , 218–228.
21. Somku i,J., Moln
´
a ,O.R., G
´
ad,A. and Smelle ,L. (2021) P essu e
pe u ba ion s udies o noncanonical i al nucleic acid s uc u es.
Biology , 10 , 1173.
22. Moln
´
a ,O.R., V
´
egh,A., Somku i,J. and Smelle ,L. (2021)
Cha ac e iza ion o a G-quad uplex om hepa i is B i us and i s
s abiliza ion by binding TMPyP4, BRACO19 and PhenDC3. Sci.
Rep. , 11 , 23243.
23. Sun,J., Wu,G., Pas o ,F., Rahman,N., Wang,W.-H., Zhang,Z.,
Me le,P., Hui,L., Sal e i,A., Du an el,D. e al. (2022) RNA helicase
DDX5 enables STAT1 mRNA ansla ion and in e e on signalling in
hepa i is B i us eplica ing hepa ocy es. Gu , 71 , 991–1005.
24. Nu k,S ., Ko en,S ., Rhie,A., Rau iainen,M., Bzikadze,A.V.,
Mikheenko,A., Vollge ,M.R., Al emose,N., U alsky,L.,
Ge shman,A. e al. (2022) The comple e sequence o a human
genome. Science , 376 , 44–53.
25. Ok onechnik o ,K., Goloso a,O., Fu so ,M. and Team,U. (2012)
Unip o UGENE: a uni ied bioin o ma ics oolki . Bioin o ma ics , 28 ,
1166–1167.
26. C ooks,G .E., Hon,G ., Chandonia,J.-M. and B enne ,S.E. (2004)
WebLogo: a sequence logo gene a o . Genome Res. , 14 , 1188–1190.
27. Ha kins,K.M. and S one,A.C. (2015) Ancien pa hogen genomics:
insigh s in o iming and adap a ion. J. Hum. E ol. , 79 , 137–149.
28. de-Dios,T., Scheib,C.L. and Houldc o ,C.J. (2023) An adagio o
i uses, played ou on ancien DNA. Genome Biol. E ol. , 15 , e ad047.
29. Taubenbe ge ,J.K., Bal imo e,D., Dohe y,P.C., Ma kel,H.,
Mo ens,D.M., Webs e ,R.G. and Wilson,I.A. (2012) Recons uc ion
o he 1918 in luenza i us: unexpec ed ewa ds om he pas . Mbio ,
3 , e00201-12.
30. Sudhan,S .S . and Sha ma,P. (2020) Human i uses: eme gence and
e olu ion. Eme g. Reeme g. Vi al Pa hog. , 2020 , 53–68.
31. I ing-Pease,E.K., Muk upa ela,R., Dannemann,M. and Racimo,F.
(2021) Quan i a i e human paleo gene ics: w ha can ancien DN A ell
us abou complex ai e olu ion? F on . Gene . , 12 , 703541.
32. Sch
¨
adle ,S. and Hild ,E. (2009) HBV li e cycle: en y and
mo phogenesis. Vi uses , 1 , 185–209.
33. Boh
´
alo
´
a,N., Can a a,A., Ba as,M., Kau a,P.,
ˇ
S
ˇ
as n´y,J., Pe
ˇ
cinka,P.,
Foj a,M. and B
´
azda,V. (2021) T acing dsDNA i us-hos
coe olu ion h ough co ela ion o hei G-quad uplex- o ming
sequences. In . J. Mol. Sci. , 22 , 3433.
34. Ba as,M., B
´
azda,V., Boh
´
alo
´
a,N., Can a a,A., Voln
´
a,A.,
S achu o
´
a,T., Malacho
´
a,K., J agelsk
´
a,E.B., P o ubiako
´
a,O.,
ˇ
Ce e
ˇ
n,J. e al. (2020) In-dep h bioin o ma ic analyses o nido i ales
including human SARS-CoV-2, SARS-CoV, MERS-CoV i uses
sugges impo an oles o non-canonical nucleic acid s uc u es in
hei li ecycles. F on . Mic obiol. , 11 , 1583.
35. La igne,M., Helynck,O., Rigole ,P., Boud ia-Souilah,R.,
No wako wski,M., Ba on,B., B ¨ul
´
e,S., Hoos,S., Raynal,B., Gui a ,L.
e al. (2021) SARS-CoV-2 Nsp3 unique domain SUD in e ac s wi h
guanine quad uplexes and G4-ligands inhibi his in e ac ion. Nucleic
Acids Res. , 49 , 7695–7712.
36. Ely,B. (2021) Genomic GC con en d i s downwa d in mos bac e ial
genomes. PLoS One , 16 , e0244163.
37.
ˇ
Sma da,P ., Bu e
ˇ
s,P ., Ho o
´
a,L., Lei ch,I.J., Mucina,L., Pacini,E.,
Tich´y,L., G ulich,V. and Ro eklo
´
a,O. (2014) Ecological and
e olu iona y signi icance o genomic GC con en di e si y in
monoco s. P oc. Na l. Acad. Sci. U.S.A. , 111 , E4096–E4102.
38. Wang,Y., Mao,J.-M., Wang,G.-D., Luo,Z.-P., Yang,L., Yao,Q. and
Chen,K.-P. (2020) Human SARS-CoV-2 has e ol ed o educe CG
dinucleo ide in i s open eading ames. Sci. Rep. , 10 , 12331.
39. Ma y
´
a
ˇ
sek,R. and Ko a
ˇ
´
ık,A. (2020) Mu a ion pa e ns o human
SARS-CoV-2 and ba RaTG13 co ona i us genomes a e s ongly
biased owa ds C > U ansi ions, indica ing apid e olu ion in hei
Hos s. Genes (Basel) , 11 , 761.
40. Goswami,P., Ba as,M., Lexa,M., Boh
´
alo
´
a,N., Voln
´
a,A.,
ˇ
Ce e
ˇ
n,J.,
ˇ
Ce e
ˇ
no
´
a,V., Pe
ˇ
cinka,P.,
ˇ
Spunda,V., Foj a,M. e al. (2021)
SARS-CoV-2 ho -spo mu a ions a e signi ican ly en iched wi hin
in e ed epea s and CpG island loci. B ie Bioin o m , 22 , 1338–1345.
41. B
´
azda,V., Po ubiako
´
a,O., Can a a,A., Boh
´
alo
´
a,N., Cou al,J.,
Ba as,M., Foj a,M. and Me gny,J.-L. (2021) G-quad uplexes in
H1N1 in luenza genomes. BMC Genomics [Elec onic Resou ce] , 22 ,
77.
42. Boh
´
alo
´
a,N., Can a a,A., Ba as,M., Kau a,P.,
ˇ
S
ˇ
as n´y,J., Pe
ˇ
cinka,P.,
Foj a,M., Me gny,J.-L. and B
´
azda,V. (2021) Analyses o i al
genomes o G-quad uplex o ming sequences e eal hei co ela ion
wi h he ype o in ec ion. Biochimie , 186 , 13–27.
43. B own,J.C. (2007) High G+C con en o he pes simplex i us DNA:
p oposed ole in p o ec ion agains e o ansposon inse ion. Open
Biochem. J. , 1 , 33–42.
44. Vinog ado ,A.E. and Ana skaya,O.V. (2017) DNA helix: he
impo ance o being AT- ich. Mamm. Genome , 28 , 455–464.
45. Cal ignac-Spence ,S., D¨ux,A., Goga en,J.F. and Pa ono,L.V. (2021)
Chap e Two - Molecula a cheology o human i uses. In:
Kielian,M., Me enlei e ,T.C. and Roossinck,M.J. (eds.) Ad ances in
Vi us Resea ch . Academic P ess, Vol. 111 , pp. 31–61.
C
The Au ho (s) 2023. Published by Ox o d Uni e si y P ess on behal o Nucleic Acids Resea ch.
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p: // c ea i ecommons.o g / licenses / by / 4.0 / ), which
pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed.
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024