7198–7204 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 Published online 3 July 2023
h ps://doi.o g/10.1093/na /gkad556
G-quad uplexes in he e olu ion o hepa i is B i us
V
´
acla B
´
azda
1 ,
*
, Michaela Dob o oln
´
a
1 , 2
, Na
´
alia Boh
´
alo
´
a
1 and Jean-Louis Me gny
1 , 3 ,
*
1
Ins i u e o Biophysics o he Czech Academy o Sciences, B no, Czech Republic,
2
Facul y o Chemis y, B no
Uni e si y o Technology, Pu ky
ˇ
no a 118, 612 00 B no, Czech Republic and
3
Labo a oi e d’Op ique e Biosciences
(LOB), Ecole Poly echnique, CNRS, INSERM, Ins i u Poly echnique de Pa is, 91120 Palaiseau, F ance
Recei ed Ma ch 30, 2023; Re ised May 23, 2023; Edi o ial Decision June 14, 2023; Accep ed June 19, 2023
ABSTRACT
Hepa i is B i us (HBV) is one o he mos dange ous
human pa hogenic i uses ound in all co ne s o he
wo ld. Recen sequencing o ancien HBV i uses e-
ealed ha hese i uses ha e accompanied human-
i y o se e al millenia. As G-quad uplexes a e con-
side ed o be po en ial he apeu ic a ge s in i ol-
ogy, we examined G-quad uplex- o ming sequences
(PQS) in mode n and ancien HBV genomes. Ou
anal yses sho wed he p esence o PQS in all 232
es ed HBV genomes, wi h a o al numbe o 1258 mo-
i s and an a e a ge equenc y o 1.69 PQS pe kbp.
No ably, he PQS wi h he highes G4Hun e sco e in
he e e ence genome is he mos highly conse ed.
In e es ingly, he densi y o PQS mo i s is lowe in
ancien HBV genomes han in hei mode n coun e -
pa s (1.5 and 1.9 / kb, espec i ely). This mode n e-
quency o 1.90 is e y close o he PQS equency o
he human genome (1.93) using iden ical pa ame e s.
This indica es ha he PQS con en in HBV inc eased
o e ime o become close o he PQS equency in
he human genome. No s a is ically signi ican di -
e ences we e ound be ween PQS densi ies in HBV
lineages ound in di e en con inen s. These esul s,
which cons i u e he i s paleogenomics analysis o
G4 p opensi y, a e in ag eemen wi h ou hypo he-
sis ha , o i uses causing ch onic in ec ions, hei
PQS equencies end o con e ge e olu iona il y wi h
hose o hei hos s, as a kind o ‘gene ic camou lage’
o bo h hijack hos cell ansc ip ional egula o y sys-
ems and o a oid ecogni ion as o eign ma e ial.
GRAPHICAL ABSTRACT
INTRODUCTION
Hepa i is B i us (HBV) belongs o he genus O hohep-
adna i us; i s genome is cons i u ed o double-s anded
DNA. This i us causes Hepa i is B , a highly con a-
gious, po en ially a al disease ha a ec s an es ima ed
257 million people wo ldwide, esul ing in an es ima ed
820 000 dea hs e e y yea ( h ps://www.who.in /news- oom/
ac - shee s/de ail/hepa i is- b ). This i us is pa icula ly se-
e e, as a pp oxima el y one in i e ca ie s die om ci ho-
sis and / o de elop hepa ocellula ca cinoma. HBV is ans-
mi ed p ima ily h ough blood and body luids and he
incuba ion pe iod is a iable, usually be ween 30 and 180
days. Du ing eplica ion, HBV DNA o ms a minich omo-
some in he nucleus o in ec ed hepa ocy es ( 1 , 2 ) and i s
genome is eplica ed h ough a p ocess o e e se ansc ip-
ion o he key in e media e p e-genomic RNA in hepa o-
cy es, which is also an mRNA empla e o he HBV p o-
eins ( 3 , 4 ). Hepadna i uses in ec ing o he hos s ha e e-
cen ly been iden i ied, including ba s , ogs , liza ds , ish, and
he capuchin monkey ( 5–8 ). Analyses o ancien genomes
ha e e ealed ha he mos ecen common ances o o all
HBV lineages is es ima ed o ha e exis ed be ween ∼20 000
*
To whom co espondence should be add essed. Tel: +420 541517231; Email: acla @ibp.cz
Co espondence may also be add essed o Jean-Louis Me gny. Tel: +33 766290967; Email: jean-louis.me gny@poly echnique.edu
C
The Au ho (s) 2023. Published by Ox o d Uni e si y P ess on behal o Nucleic Acids Resea ch.
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p: // c ea i ecommons.o g / licenses / by / 4.0 / ), which
pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed.
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7199
and 12 000 yea s ago, and he i us was ound o be p esen
in Eu opean and Sou h Ame ican hun e -ga he e s du ing
he ea ly Holocene pe iod ( 9 ).
A la ge amoun o li e a u e is de o ed o he occu ence
o G-quad uplexes in i uses, especially o he possibili y o
using hese s uc u es in he apy. Comp ehensi e bioin o -
ma ics analyses ha e aced pu a i e G4- o ming sequences
in he genome o almos all human i uses, showing ha
hei dis ibu ion and p esence a e highly conse ed. The e-
o e, hese DNA o RNA s uc u es can be a sui able a -
ge o a ge ed he apy. Some G-quad uplex ligands ha e
been shown o ha e an i i al ac i i y, o example, agains
HIV ( 10 ), he pes simplex i us I (HSV-1) ( 11 ), SARS-CoV-
2 ( 12 ) and o he s. A comp ehensi e analysis o all sequenced
i uses ha ha e a la en phase in hei li e cycle showed
ha hei G-quad uplex con en is co ela ed wi h ha o
he hos ( 13 ). In con as , i uses causing acu e in ec ions
wi hou a la en phase end o elimina e G-quad uplex se-
quences, as hey can become oadblocks du ing eplica ion,
ansc ip ion and / o e e se ansc ip ion ( 14 ).
Mo e speci ically, G4s ha e been ound o be ele an
in HBV in ec ion, bo h a he DN A and RN A le el ( 15 ).
Chak abo y and Ghosh epo ed ha an RNA sequence
p esen in HBV RNA exhibi ed a sequence-independen
ans-ac ing nuclease ac i i y, and ha his sequence adop s
a G4 con o ma ion ( 16 ). Biswas e al. analyzed a G4 p one
mo i ( GGGAGTGGGAGCATTCGGGCCAGGG ) ha is highly
conse ed only in HBV geno ype B, and was shown o
adop a hyb id s uc u e ( 17 ). In e es ingl y, m u a ions dis-
up ing his G-quad uplex in HBV geno ype B cons uc s
we e associa ed wi h impai ed i ion sec e ion. The au ho s
p oposed ha his G4 media es enhancemen o ansc ip-
ion and i ion sec e ion in his HBV geno ype. In a la e e-
iew, hey no e ha among i uses con aining a G4 in hei
genome, hose associa ed wi h cance a e o e - ep esen ed,
including HBV ( 18 ).
In con as , some G4s end o be conse ed in all geno-
ypes, as epo ed by Meie -S ephenson e al ( 19 ) o a
DNA sequence ound in he p e-co e p omo e egion
( CTGGGAGGAGCTGGGGGAGGAGA ). They demons a ed a
ole o his quad uplex in i al eplica ion by compa ing
he wild- ype mo i o non-G4- o ming mu an s in i o .
Fleming e al iden i ied conse ed po en ial G4 sequences
in se e al i al genomes ele an o human heal h, and
showed ha hese mo i s can p o ide a ame wo k o N6-
me hyladenosine (m6A) ins alla ion wi hin he loops o
RNA G4 sequences ound in se e al i uses, including HBV
( 20 ). Somku i e al. de e mined olume changes in h ee
HBV G4 s uc u es using biophysical app oaches in i o .
They in es iga ed h ee DNA sequences: GGCTGGGGCTTG-
GTCATGGGCCATCAG , GGGAGTGGGAGCATTCGGGCCAGGG
and TTGGGTGGCTTTGGGGCATGGAC ( 21 ). The same g oup
in es iga ed one o hese sequences in mo e de ails (HepB;
GGCTGGGGCTTGGTCATGGGCCATCAG , ound in he coding
egion o he polyme ase p o ein), and analyzed i s in e -
ac ion wi h G4 ligands in i o ( 22 ). Finally, Sun e al. e-
cen ly in es iga ed he ole o cellula G4s (mo i s ound in
he hos cell genome) in HBV in ec ion, demons a ing ha
he DDX5 helicase, known o be capable o esol ing RNA
G4 s uc u es, is a key egula o o he in e e on (IFN)
esponse agains his i us. DDX5 down egula ion is ob-
se ed du ing HBV eplica ion and in poo p ognosis HBV-
ela ed hepa ocellula ca cinoma (HCC) ( 23 ). All o hese
esul s poin ou he links be ween i al (o hos ) DNA and
RNA G4s and HBV.
In his pape , we ha e analyzed 232 HBV genomes om
samples co e ing a mo e han 10- housand-yea his o y o
he p esence o G-quad uplex o ming sequences. Ou e-
sul s show an e olu iona y shi o an inc eased numbe o
G-quad uplex es in ecen HBV i uses, poin ing o he im-
po ance o G-quad uplexes in he HBV li e-cycle in human
li e cells.
MATERIALS AND METHODS
Genomes
232 HBV alignmen s we e downloaded om he supple-
men a y ma e ials a (9). Sequences we e ob ained o 122
mode n geno ypes and di e en g oups o 110 ancien
s ains, di ided in o g oups based on mPTP classi ica ion.
As he e e ence genome, we ook NC 003977.2 and, o
analyses o phylogene ically ela ed i uses wi h hos s o he
han human, we il e ed e e ence genomes om O hohep-
adna i uses. In o al, we downloaded 21 addi ional HBVs
ha ing a non-human hos , in ec ing bi ds (7 genomes), ba s
(5 genomes), ish (2 genomes), o he Mammals (i.e. nei he
human no ba s; 6 genomes) and amphibians (1 genome, Ti-
be an og hepa i is B i us). HBV G4 con en s we e com-
pa ed o he gapless human genome, he new elome e- o-
elome e assembly o he human genome ( 24 ), which was
downloaded om NCBI (T2T-CHM13 2.0).
G4Hun e analyses
All sequences we e analyzed using G4Hun e
( h p://bioin o ma ics .ibp .cz ) o iden i y PQS sequences.
G4Hun e ’s de aul pa ame e s we e used (25 nucleo ides
o window size and 1.2 o h eshold). These se ings ha e
p e iously been shown o iden i y expe imen ally- alida ed
quad uplex s uc u es. The lis o all o ganisms es ed
and he esul s o he analyses we e downloaded om he
supplemen a y ma e ials a ( 9 ).
S a is ical e alua ion
Da a wi h G4Hun e esul s we e me ged in an Excel ile o
s a is ical e alua ion. G4Hun e esul s , leng hs , and GC
con en o analyzed sequences a e accessible in Supplemen-
a y ma e ial 01. A sca e plo was gene a ed in G aph-
Pad P ism ( 8.0.1), Violin plo s we e cons uc ed in R (
4.2.0) wi h ggplo 2. S a is ical signi icance was es ed using
S uden ’s T- es . No mali y o da a was de e mined using
Shapi o-Wilk es .
Cons uc ion o LOGO sequence
All sequences o ancien and mode n HBV genomes we e
uploaded in o UGENE so wa e ( 25 ) and he loca ion o
PQS sequences we e ex ac ed using Clus alW alignmen .
LOGO sequence was gene a ed in aligned sequences and
WebLogo 3 ool ( 26 ).
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
7200 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14
Table 1. S a is ic da a o G4Hun e analyses o HBV i uses. n Seq (numbe o s ains), Leng h (leng h o he sequence, n ), GC % (a e age GC con en ),
PQS n ( o al numbe o p edic ed PQS wi h a G4Hun e sco e o 1.2 o mo e), Mean PQS (a e age numbe o p edic ed PQS), Min PQS (lowes equency
o p edic ed PQS), Max PQS (highes equency o p edic ed PQS), PQS pe 1000 GC (PQS equency pe 1000 GC)
G oups Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC
All HBV 232 3214 45.8 1258 1.69 0.61 4.04 3.66
G oups Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC
ancien 122 3217 43.1 587 1.50 0.61 2.83 3.46
mode n 110 3210 48.8 671 1.90 0.94 4.04 3.88
H. sapiens T2T 1* 3.05 ×10
9 41.6 5.45 10
6 1.93 - - 4.72
Subg oups** Seq n Leng h n GC % PQS n Mean PQS Min PQS Max PQS PQS pe GC
A 21 3227 47.0 105 1.55 2.00 8.00 3.30
B 14 3222 49.0 120 2.66 1.55 4.04 5.43
C 16 3222 49.1 96 1.86 1.55 3.10 3.79
D 57 3190 47.7 317 1.74 0.94 2.83 3.66
E 3 3218 48.4 25 2.59 2.17 2.80 5.35
F 13 3222 48.8 81 1.93 1.24 2.48 3.96
G 3 3255 48.2 12 1.23 1.23 1.23 2.55
H 5 3222 48.9 39 2.42 1.86 3.11 4.94
I 3 3219 48.8 20 2.07 1.55 2.49 4.24
J 1 3187 49.0 4 1.26 1.26 1.26 2.56
Ancien Ame ican 4 3214 38.4 15 1.17 0.93 1.55 3.12
Ea ly Ana olian a me 1 3189 33.7 2 0.63 0.63 0.63 1.86
Mesoli hic 8 3192 45.8 33 1.29 0.63 2.20 3.20
WENBA 65 3229 42.1 288 1.37 0.61 2.20 3.27
o u 2 3191 49.1 8 1.25 1.25 1.25 2.56
gbn 3 3194 49.1 20 2.09 1.57 2.51 4.25
czp 3 3191 48.3 18 1.88 1.25 2.51 3.89
o he 10 3211 45.2 55 1.71 0.62 4.04 3.81
*Comple e elome e- o- elome e human genome (22 + X + Y ch omosomes)
**The g oups we e de ined by Koche e al. ( 10 ) and is based on he mul i- a e Poisson T ee P ocesses (mPTP) as a esul s o gene ic clus e s numbe
conside ing a phylogene ic inpu ee ( 26 ).
RESULTS
We analyzed he p esence o PQS in 232 HBV genomes
(122 ancien and 110 ecen ) using G4Hun e . All human
HBV genomes a e simila in leng h, a ying om 3180 o
3300 bp. Compa isons o ancien and ecen samples show
a sligh and non-signi ican change in a e age leng h om
3217 o 3210 bp. On he o he hand, hese genomes a y in
GC con en , and we ound he p esence o G-quad uplex-
o ming sequences in all HBV genomes in he da ase . In
o al, we ound 1258 PQS wi h a mean equency o 1.69
PQS pe kbp (Table 1 ). The mean equency o PQS in
ancien genomes is 1.50 / kb, compa ed wi h he mean e-
quency in mode n HBV genomes o 1.90. Fo compa ison
we also analyzed he ne wly pub lished human gapless assem-
bly. The PQS equency in he human genome is 1.93, which
is almos iden ical o he a e age PQS equency o mode n
HBV genomes (Table 1 ). Al hough he e is no signi ican
change in he leng h o ancien and mode n HBV genomes,
compa ison o PQS densi y shows ha he mode n i uses
a e subs an ially iche in PQS (Figu e 1 ) ( P - alue = 3.6e-
08). The mode n HBV genomes no only ha e a highe PQS
equency, bu also ha e a highe GC con en . To e alua e i
he change in PQS equency is s a is ically signi ican a e
aking in o accoun GC con en , we ecalcula ed he PQS
equency acco ding o GC con en (Table 1 , las column).
E en a e his co ec ion, he PQS equency / GC con en
is highe in mode n han in ancien HBV genomes (3.88 e -
sus 3.46 pe housand GC o he mode n and ancien HBV
genomes, espec i ely: P - alue = 1.8e-03).
We hen u he di ided genomes acco ding o di e -
en geno ypes. While mos cu en geno ypes ha e an a -
Figu e 1. Compa ison o PQS equencies in ancien and mode n HBV
genomes.
e age PQS equency highe han 2 PQS / kb, ou o he
i e ancien genomes ha e PQS equencies lowe han his
alue. The highes PQS equencies we e ound in mod-
e n geno ypes B, E and H, he lowes in ancien Ame ican,
Mesoli hic and o he ancien geno ypes (Figu e 2 ).
We p esen he PQS equency pe kb ( o all mo i s wi h
a G4Hun e sco e > 1.2) and he leng h o he genome as a
unc ion o ime o each ancien sequence (Supplemen a y
ma e ial 02). As men ioned be o e, he longes sequence
only di e s om he sho es by 120 bp, and he leng h o
he genome does no change signi ican ly o e ime (Pea -
son = −0.1387, P ( wo- ailed) = 1.6e-01). Unlike genome
leng h, PQS equencies pe kbp we e ound o inc ease o e
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7201
Figu e 2. PQS equency o HBV subg oups di ided in o ancien and mode n geno ypes. Ancien Ame ican 1, 2, 3 ca ego ies we e me ged. Mesoli hic 1
and 2 geno ypes we e me ged.
Table 2. Posi ion o G4 sequences in he human HBV e e ence genome (NC 003977.2) and i s loca ion conse a ion. A nega i e sco e co esponds o a
C- ich sequence, meaning ha i is i s complemen a y s and which will be G4-p one
Posi ion S and Sequence G4H sco e Fea u e / loca ion G4 conse a ion
316 −CCCC AA CC T CC AAT C A C T C A CC −1.41 S, DNA polyme ase
N- e minal domain 84.5%
737 + GG AT G AT G T GG TATT GGGGG +1.50 S, DNA polyme ase
C- e minal domain 70.3%
1133 −CC TGAA CC TTTA CCCC GTTG CCC −1.30 P, DNA polyme ase
C- e minal domain 19.4%
1722 + GGG A GG AGTT GGGGG A GG A G ATTA GG +1.65 T ansac i a ion p o ein X 92.2%
1887 + GGG T GG CTTT GGGG CAT GG +1.63 C , Hepa i is co e p o ein 40.5%
ime (Pea son = 0.2903, P ( wo- ailed) = 2.8e-03; Supple-
men a y ma e ial 02).
The whole genome o HBV is ansc ibed in o one long
p e-genomic RN A (mRN A-pgRN A) ha encodes all HBV
p o eins, and pgRNA also ac s as a empla e o e e se
ansc ip ion. Analysis o PQS localiza ion in he HBV e -
e ence genome (NC 003977.2) shows ha all PQS a e lo-
ca ed in he icini y o gene egions, which is no su p ising
conside ing ha he HBV genome is small and he en i e
genome is used e y e ec i ely o p oduce he ew p o eins
necessa y o i s unc ion, such as DN A pol yme ase, ans-
ac i a ion and capsid p o eins (Table 2 ).
The PQS wi h he highes G4Hun e sco e in he HBV
e e ence genome is loca ed su ounding posi ion 1722, in
he egion ha codes o ansac i a ion p o ein X. In e -
es ingly, a PQS is p esen in almos all HBV genomes a
his loca ion (214 o 232; o 92.2% o HBV genomes an-
alyzed), making i he mos conse ed mo i be ween all
PQS in he e e ence HBV genome. Conse a ion implies
ha his sequence posi ion has been main ained by selec-
i e p essu e. Compa ison o he LOGO sequence o his
loca ion in mode n and ancien HBV genomes (Figu e 3 )
demons a es ha a G- ich mo i is p ese ed in all s ains.
Ne e heless, his ‘G- ichness’ is e en mo e s iking in mod-
e n compa ed o ancien s ains, wi h Gs becoming p edom-
inan a posi ions 8 and 11 (Figu e 3 , a ows), while ancien
Figu e 3. L OGO ep esen a ion o he consensus mo i ound a ound po-
si ion 1722 (acco ding o he e e ence HBV genome) in mode n and an-
cien HBV genomes. Guanine nucleo ides p edomina e in mode n s ains
a posi ions 8 and 11 (a ows) while in ancien HBV genomes o he nu-
cleo ides a e p esen (T and A in posi ion 8). A less s iking G-en ichmen
is also ound a posi ions 14 and 15. As a consequence, while bo h con-
sensus mo i s a e compa ible wi h G4 o ma ion, he mode n sequence is
mo e a o able han he ancien one (G4hun e sco es o 2.11 and 1.83,
espec i ely).
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
7202 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14
Table 3. G4Hun e sco e in HBV i uses, g ouped by con inen and age
G oups Seq Leng h Mean PQS Min PQS Max PQS
All 232 3214 1.69 0.61 4.04
Ancien Seq Leng h Mean PQS Min PQS Max PQS
Aus alia 0 - - - -
Ame ica 7 3217 1.33 0.93 3.11
A ica 1 3229 0.62 - -
Asia 34 3219 1.54 0.61 2.51
Eu ope 82 3217 1.50 0.61 2.83
Mode n Seq Leng h Mean PQS Min PQS Max PQS
Aus alia 14 3215 1.93 0.94 3.10
Ame ica 31 3216 1.96 0.94 3.11
A ica 15 3205 1.83 1.24 2.80
Asia 110 3206 1.93 0.94 4.04
Eu ope 12 3209 1.66 1.23 2.48
HBV genomes mo e o en exhibi T / A o A a hese posi-
ions. A highe –– nea 100% –– p e alence o Gs is also is-
ible a posi ions 14 and 15 (Figu e 3 ) in mode n genomes.
O e all, while bo h ancien and mode n consensus mo i s
a e G4-p one, he mode n sequence is mo e a o able, as
shown by he highe G4Hun e sco e.
We also di ided HBV genomes acco ding o he geo-
g aphic place o sampling. Mos o he ancien HBV sam-
ples we e ound in Eu ope, only one in A ica and none in
Aus alia. None heless, ancien genomes ha e a lowe PQS
equency han mode n HBV genomes, ega dless o hei
con inen o o igin (Table 3 ; also see G aphical Abs ac ).
DISCUSSION
Pa hogens e ol e in esponse o human biological changes
alongside sociocul u al and echnological de elopmen s
( 27 ). Ancien i al genomes p o ide in o ma ion on he
e olu ion o i uses o e bo h ime and space and p o-
ide insigh in o he changes ha may ha e occu ed in
i ulence and ansmissibili y ( 28 ). Cu en ad anced ech-
niques o isola ion o nucleic acids and sequencing ha e al-
lowed paleogenomic o ‘a cheo i ology’ in es iga ions. The
in amous 1918 ‘Spanish’ in luenza pandemic was he sou ce
o he i s ancien pa hogen genome ( 29 ) and a cheo i ol-
ogy has been g owing apidly since hen. Conside ing ha
wo- hi ds o all human pa hogens a e i uses ( 30 ), paleo-
gene ics o i al genomes p o ides an in e es ing iewpoin
on human his o y ( 31 ).
While se e al ancien i al genomes a e a ailable, he
mos comp ehensi e da ase deals wi h ancien HBV
genomes ( 28 ). Wi h mo e han 200 million people su e ing
om ch onic HBV in ec ion, HBV can be conside ed as a
common i us. HBV has a li e-cycle ha equi es i s double-
s anded DNA genome o each he hos cell nucleus ( 32 ),
in con as wi h RNA i uses such as in luenza o SARS-
CoV-2 ha cause only acu e in ec ions and con ain RNA
genomes ha can be eplica ed and ansla ed in he cy o-
plasm. Vi uses causing acu e in ec ions end o ha e a low
PQS equency, while G4s end o be ubiqui ous in mos o -
ganisms.
G4s in he HBV genome a e impo an s uc u es ha
egula e ansc ip ion and i ion sec e ion in HBV geno-
ype B ( 15 , 17 ). In his epo , we analyzed he p esence o
PQS in mul iple HBV genomes, om ancien o cu en
HBV s ains. Impo an ly, we ound ha PQS equency is
highe in ecen compa ed o ancien s ains. I was shown
p e iously ha he G4 equency o dsDNA i uses co e-
la es wi h he PQS equency o he hos , shown o dsDNA
i uses in ec ing A chaea, Bac e ia and Euka yo a ( 33 ). In
ag eemen wi h hose da a, we ound ha he densi y o G4
mo i s in mode n HBV s ains (dsDNA i uses ha also
expe ience a la en phase) ends o con e ge o he o e all
G4 densi y o he human genome. We p opose ha mu a-
ions which led o ‘PQS in eg a ion’ (new G4 mo i s wi hin
he HBV genome) we e e olu iona y p e e ed and ixed in
HBV pa hogenic s ains du ing e olu ion. I seems ha he
opposi e p ocess may occu in i uses causing acu e in ec-
ions, as ound o SARS-CoV-2, whe e he PQS equency
is ex emely low compa ed o he PQS equencies o o he
co ona i uses ( 34 , 35 ).
Vi uses ha e highly a iable genomes and a e p one o
mu a ions, in con as o cellula and especially mul icellu-
la o ganisms, as desc ibed epea edly. This is especially ue
o RN A i uses, w he e he m u a ion a e is se e al o de s
o magni ude highe han in DNA-based genomes. A d a-
ma ic dec ease in GC con en has been desc ibed o se e al
bac e ial species ( 36 ) and o some plan species wi h holo-
cen ic ch omosomes ( 37 ). Simila ly, a apid dec ease in GC
con en o e ime was ound o some i uses causing acu e
in ec ions (and wi hou la ency connec ed o nuclea local-
iza ion), including Nido i ales, in luenza genomes and he
con empo a y SARS-CoV-2 ou b eak ( 38 ). Compa ison o
SARS-CoV-2 genomes showed a s ong p e e ence o mu-
a ions in GC islands and C > U ansi ions, leading o a
dec ease in G4 p opensi y ( 39 , 40 ). As a consequence, he
PQS equency o hese i uses causing only acu e in ec-
ions is gene ally e y low (0.03 o SARS-2 ( 35 ), 0.56 o
in luenza H1N1 genomes ( 41 )).
The opposi e end is ue o HBV, which exhibi s a la-
en s a e and main ains i s genome in he nucleus: mode n
HBV genomes ha e a signi ican ly highe PQS equency
compa ed o ancien HBV genomes. Ou esul s a e in line
wi h a b oad s udy compa ing PQS equencies in i uses
wi h p edominan ly pe sis en o acu e ypes o in ec ion
( 42 ). Acco ding o ha s udy, i uses causing pe sis en in-
ec ion a e en iched in PQS compa ed o acu ely in ec ious
i uses. Impo an ly, his obse a ion is also alid wi hin
i uses causing hepa i is: HAV (hepa i is A i us - caus-
ing acu e in ec ion) ha e a low PQS equency, while he
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
Nucleic Acids Resea ch, 2023, Vol. 51, No. 14 7203
Hepadna i idae ha cause ch onic in ec ions ( o which hu-
man HBV belongs) ha e a signi ican ly highe PQS e-
quency ( 42 ).
The p esence o G4s depends on guanine con en in he
genome, and one o he possible ad an ages o a GC- ich
genome is o p o ide addi ional gene egula ion oppo u-
ni ies. In his espec , G4s ha e been shown o be impo -
an o ansc ip ion in highe o ganisms. An inc ease in
GC con en has also been documen ed in plan species ha
can g ow in seasonally cold clima es, possibly indica ing
an ad an age o GC- ich DNA du ing cell eezing, and
he genomic adap a ions associa ed wi h changing GC con-
en a e sugges ed o g ass-domina ed biomes du ing he
Te ia y pe iod ( 37 ). G4s a e o en o e ep esen ed in he
p omo e egions o highe euka yo es and ha e also been
demons a ed o con ibu e o di ec ed genome edi ing in
nema odes. Fo i uses expe iencing a la en phase in he
nucleus, ha ing a simila genome o ganiza ion as he hos
is ad an ageous o bo h a oid ecogni ion as unusual ( o -
eign) DNA, and o hijack he hos egula o y machine y. In
addi ion, he high GC con en o HSV DNA is sugges ed o
ac as a p o ec i e ea u e agains e o ansposon inse ion
( 43 ).
As AT- ich egions in humans a e mos ly associa ed wi h
condensed ch oma in ( 44 ), he shi o GC- ich i uses
could be impo an o i uses wi h la en phase o ha e a
be e chance o being ac i e in he u u e. HBV has a la-
en pe iod, he e o e, he e olu iona y p essu e o inc ease
GC con en and PQS p esence could be e olu iona y a-
o ed. In his model, he o iginal (non-human) p e-HBV
hos could ha e had a lowe G4 equency – adap a ion o
he human hos may ha e led o a hos -pa hogen PQS con-
e gence and a concomi an inc ease in G4 densi y in he
i us.
CONCLUSION
We pe o med he i s paleogenomic analysis o G4
p opensi y, applied he e o he Hepa i is B i us. We ound
ha he densi y o PQS mo i s inc eased o e ime, as i
is highe in mode n han ancien HBV genomes. The e-
quency in mode n i uses is now e y close o ha o he
human genome. This s udy should pa e he way o he
paleo genomics anal ysis o G4 sequences (and o he sim-
ila mo i s) in o he pa hogens. Un o una ely, his o ical
in o ma ion abou i uses genomes is only a ely a ail-
able: A cheo i ology is a nascen ield ( 45 ), which aces
he same obs acles as mode n genomics, bu wi h he ad-
di ional p oblem o analyzing pa ially deg aded DNA. As
no ed by he au ho s ( 45 ), HBV is ‘an excellen a ge o
he eco e y o ancien sequences due o i s ela i ely s ab le,
pa ially double-s anded ci cula DNA genome, i s high
p e alence in he human popula ion, and p olonged high
i emia du ing ch onic in ec ion’. Double-s anded DNA
is gene ally be e p ese ed han single-s anded DNA o
RNA. This may explain why HBV is cu en ly he only in-
s ance in which sequence da a a e a ailable o o e 100 an-
cien i uses. In o he si ua ions, such as a iola i us, only
a ew ancien genomes a e a ailable ( he e a e only 4, 10
and 11 HSV-1, Pa o i uses and VA RV ancien sequences
a ailab le, espec i el y ( 28 ). Anal yses o HIV-1 o he 1918
in luenza i uses is possible, bu only o e a limi ed pe iod
o ime.
A cheo i ology will be use ul o iden i y polymo phisms
impo an o human adap a ion o pa hogens, and ice-
e sa du ing he complex i us-hos ela ionships c ucial o
con inued i al p e alence ( 28 ). Mo e paleogenomic da a
will be needed o es he hypo hesis ha , o i uses causing
ch onic in ec ions, hei PQS equencies end o con e ge
e olu iona ily wi h hose o hei hos . In pa icula , gi en
ecen se ious ou b eaks, we hope ha he analysis o an-
cien i al pa hogens will p o ide c i ical knowledge abou
he na u e o new i al diseases.
DA T A A V AILABILITY
All da a a e a ailable in he manusc ip and supplemen a y
iles. HBV sequences we e uploaded om he supplemen-
a y ma e ials a ( 9 ).
SUPPLEMENT ARY DA T A
Supplemen a y Da a a e a ailable a NAR Online.
ACKNOWLEDGEMENTS
The au ho s hank P. Coa es o English p oo eading and
aluable commen s, and L. Gui a (LOB) o help ul dis-
cussions.
FUNDING
Agence de l’Inno a ion de D
´
e ense (AID) ia he Cen e
In e disciplinai e d’E udes pou la D
´
e ense e la S
´
ecu i
´
e
(CIEDS) [p ojec 2023 - Pa hogens]; ANR G4Access
[ANR-20-CE12-0023]; INCa G4Access g an s ( o J.L.M.);
SYMBIT p ojec [CZ.02.1.01 / 0.0 / 0.0 / 15 003 / 0000477] i-
nanced om he ERDF; Czech Science Founda ion [22-
21903S o V.B.]. Funding o open access cha ge: Academ y
o Sciences.
Con lic o in e es s a emen . None decla ed.
REFERENCES
1. Bock,C.T., Sch anz,P., Sch
¨
ode ,C.H. and Zen g a ,H. (1994)
Hepa i is B i us genome is o ganized in o nucleosomes in he
nucleus o he in ec ed cell. Vi us Genes , 8 , 215–229.
2. Newbold,J.E., Xin,H., Tencza,M., She man,G., Dean,J., Bo w den,S.
and Loca nini,S. (1995) The co alen ly closed duplex o m o he
hepadna i us genome exis s in si u as a he e ogeneous popula ion o
i al minich omosomes. J. V i ol. , 69 , 3350–3357.
3. Summe s,J. and Mason,W.S. (1982) Replica ion o he genome o a
hepa i is B–like i us by e e se ansc ip ion o an RNA
in e media e. Cell , 29 , 403–415.
4. Tiollais,P., Pou cel,C. and Dejean,A. (1985) The hepa i is B i us.
Na u e , 317 , 489–495.
5. MacDonald,D.M., Holmes,E.C., Lewis,J.C. and Simmonds,P. (2000)
De ec ion o hepa i is B i us in ec ion in wild-bo n chimpanzees
(Pan oglody es e us): phylogene ic ela ionships wi h human and
o he p ima e geno ypes. J. Vi ol. , 74 , 4253–4257.
6. D exle ,J.F., Geipel,A., K
¨
onig,A., Co man,V.M., an Riel,D.,
Leij en,L.M., B eme ,C.M., Rasche,A., Co on ail,V.M.,
Maganga,G.D. e al. (2013) Ba s ca y pa hogenic hepadna i uses
an igenically ela ed o hepa i is B i us and capable o in ec ing
human hepa ocy es. P oc. Na l. Acad. Sci. U.S.A. , 110 , 16151–16156.
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024
7204 Nucleic Acids Resea ch, 2023, Vol. 51, No. 14
7. Laube ,C., Sei z,S., Ma ei,S., Suh,A., Beck,J., He s ein,J., B
¨
o old,J.,
Salzbu ge ,W., Kade ali,L., B iggs,J.A.G. e al. (2017) Deciphe ing
he o igin and e olu ion o hepa i is B i uses by means o a amily o
non-en eloped ish i uses. Cell Hos Mic obe , 22 , 387–399.
8. de Ca alho Dominguez Souza,B.F., K
¨
onig,A., Rasche,A., de
Oli ei a Ca nei o,I., S ephan,N., Co man,V.M., Roppe ,P.L.,
Goldmann,N., Keppe ,R., M¨ulle ,S.F. e al. (2018) A no el hepa i is
B i us species disco e ed in capuchin monkeys sheds new ligh on
he e olu ion o p ima e hepadna i uses. J. Hepa ol. , 68 , 1114–1122.
9. Koche ,A., Papac,L., Ba que a,R., Key,F.M., Spy ou,M.A.,
H¨uble ,R., Roh lach,A.B., A on,F., S ahl,R., Wissgo ,A. e al.
(2021) Ten millennia o hepa i is B i us e olu ion. Science , 374 ,
182–188.
10. Pe one,R., Bu o skaya,E., Daelemans,D., Pal`u,G., Pannecouque,C.
and Rich e ,S.N. (2014) An i-HIV-1 ac i i y o he G-quad uplex
ligand BRACO-19. J. An imic ob. Chemo he . , 69 , 3248–3258.
11. F asson,I., Sold
`
a,P., Nadai,M., Tassina i,M., Scalab in,M.,
Gokhale,V., Hu ley,L.H. and Rich e ,S.N. (2022)
Quindoline-de i a i es display po en G-quad uplex-media ed
an i i al ac i i y agains he pes simple x i us 1. An i i al R es. , 208 ,
105432.
12. Zhai,L.-Y., Su,A.-M., Liu,J.-F., Zhao,J.-J., Xi,X.-G. and Hou,X.-M.
(2022) Recen ad ances in applying G-quad uplex o SARS-CoV-2
a ge ing and diagnosis: a e ie w. In . J. Biol. Mac omol. , 221 ,
1476–1490.
13. Puig Lomba di,E. and Londo
˜
no-Vallejo,A. (2020) A guide o
compu a ional me hods o G-quad uplex p edic ion. Nucleic Acids
Res. , 48 , 1–15.
14. Ruggie o,E. and Rich e ,S.N. (2020) Vi al G-quad uplexes: new
on ie s in i us pa hogenesis and an i i al he apy. Annu. Rep. Med.
Chem. , 54 , 101–131.
15. Teng,Y., Zhu,M., Chi,Y., Li,L. and Jin,Y. (2022) Can G-quad uplex
become a p omising a ge in HBV he apy? F on . Immunol. , 13 ,
1091873.
16. Chak abo y,D. and Ghosh,S. (2017) The epsilon mo i o hepa i is B
i us RNA exhibi s a po assium-dependen ibonucleoly ic ac i i y.
FEBS J. , 284 , 1184–1203.
17. Biswas,B., Kandpal,M. and Vi ekanandan,P. (2017) A G-quad uplex
mo i in an en elope gene p omo e egula es ansc ip ion and i ion
sec e ion in HBV geno ype B. Nucleic Acids Res. , 45 , 11268–11280.
18. Sa ana han,N. and Vi ekanandan,P. (2019) G-Quad uplexes: mo e
Than Jus a Kink in Mic obial Genomes. T ends Mic obiol. , 27 ,
148–163.
19. Meie -S ephenson,V., Badmalia,M.D., M ozowich,T., Lau,K.C.K.,
Schul z,S.K., Gemmill,D.L., Osiowy,C., an Ma le,G., Co in,C.S.
and Pa el,T.R. (2021) Iden i ica ion and cha ac e iza ion o a
G-quad uplex s uc u e in he p e-co e p omo e egion o hepa i is B
i us co alen ly closed ci cula DNA. J. Biol. Chem. , 296 , 100589.
20. Fleming,A.M., Nguyen,N.L.B. and Bu ows,C.J. (2019)
Colocaliza ion o m6A and G-quad uplex- o ming sequences in i al
RNA (HIV, zika, hepa i is B, and SV40) sugges s opological con ol
o adenosine N6-me hyla ion. ACS Cen . Sci. , 5 , 218–228.
21. Somku i,J., Moln
´
a ,O.R., G
´
ad,A. and Smelle ,L. (2021) P essu e
pe u ba ion s udies o noncanonical i al nucleic acid s uc u es.
Biology , 10 , 1173.
22. Moln
´
a ,O.R., V
´
egh,A., Somku i,J. and Smelle ,L. (2021)
Cha ac e iza ion o a G-quad uplex om hepa i is B i us and i s
s abiliza ion by binding TMPyP4, BRACO19 and PhenDC3. Sci.
Rep. , 11 , 23243.
23. Sun,J., Wu,G., Pas o ,F., Rahman,N., Wang,W.-H., Zhang,Z.,
Me le,P., Hui,L., Sal e i,A., Du an el,D. e al. (2022) RNA helicase
DDX5 enables STAT1 mRNA ansla ion and in e e on signalling in
hepa i is B i us eplica ing hepa ocy es. Gu , 71 , 991–1005.
24. Nu k,S ., Ko en,S ., Rhie,A., Rau iainen,M., Bzikadze,A.V.,
Mikheenko,A., Vollge ,M.R., Al emose,N., U alsky,L.,
Ge shman,A. e al. (2022) The comple e sequence o a human
genome. Science , 376 , 44–53.
25. Ok onechnik o ,K., Goloso a,O., Fu so ,M. and Team,U. (2012)
Unip o UGENE: a uni ied bioin o ma ics oolki . Bioin o ma ics , 28 ,
1166–1167.
26. C ooks,G .E., Hon,G ., Chandonia,J.-M. and B enne ,S.E. (2004)
WebLogo: a sequence logo gene a o . Genome Res. , 14 , 1188–1190.
27. Ha kins,K.M. and S one,A.C. (2015) Ancien pa hogen genomics:
insigh s in o iming and adap a ion. J. Hum. E ol. , 79 , 137–149.
28. de-Dios,T., Scheib,C.L. and Houldc o ,C.J. (2023) An adagio o
i uses, played ou on ancien DNA. Genome Biol. E ol. , 15 , e ad047.
29. Taubenbe ge ,J.K., Bal imo e,D., Dohe y,P.C., Ma kel,H.,
Mo ens,D.M., Webs e ,R.G. and Wilson,I.A. (2012) Recons uc ion
o he 1918 in luenza i us: unexpec ed ewa ds om he pas . Mbio ,
3 , e00201-12.
30. Sudhan,S .S . and Sha ma,P. (2020) Human i uses: eme gence and
e olu ion. Eme g. Reeme g. Vi al Pa hog. , 2020 , 53–68.
31. I ing-Pease,E.K., Muk upa ela,R., Dannemann,M. and Racimo,F.
(2021) Quan i a i e human paleo gene ics: w ha can ancien DN A ell
us abou complex ai e olu ion? F on . Gene . , 12 , 703541.
32. Sch
¨
adle ,S. and Hild ,E. (2009) HBV li e cycle: en y and
mo phogenesis. Vi uses , 1 , 185–209.
33. Boh
´
alo
´
a,N., Can a a,A., Ba as,M., Kau a,P.,
ˇ
S
ˇ
as n´y,J., Pe
ˇ
cinka,P.,
Foj a,M. and B
´
azda,V. (2021) T acing dsDNA i us-hos
coe olu ion h ough co ela ion o hei G-quad uplex- o ming
sequences. In . J. Mol. Sci. , 22 , 3433.
34. Ba as,M., B
´
azda,V., Boh
´
alo
´
a,N., Can a a,A., Voln
´
a,A.,
S achu o
´
a,T., Malacho
´
a,K., J agelsk
´
a,E.B., P o ubiako
´
a,O.,
ˇ
Ce e
ˇ
n,J. e al. (2020) In-dep h bioin o ma ic analyses o nido i ales
including human SARS-CoV-2, SARS-CoV, MERS-CoV i uses
sugges impo an oles o non-canonical nucleic acid s uc u es in
hei li ecycles. F on . Mic obiol. , 11 , 1583.
35. La igne,M., Helynck,O., Rigole ,P., Boud ia-Souilah,R.,
No wako wski,M., Ba on,B., B ¨ul
´
e,S., Hoos,S., Raynal,B., Gui a ,L.
e al. (2021) SARS-CoV-2 Nsp3 unique domain SUD in e ac s wi h
guanine quad uplexes and G4-ligands inhibi his in e ac ion. Nucleic
Acids Res. , 49 , 7695–7712.
36. Ely,B. (2021) Genomic GC con en d i s downwa d in mos bac e ial
genomes. PLoS One , 16 , e0244163.
37.
ˇ
Sma da,P ., Bu e
ˇ
s,P ., Ho o
´
a,L., Lei ch,I.J., Mucina,L., Pacini,E.,
Tich´y,L., G ulich,V. and Ro eklo
´
a,O. (2014) Ecological and
e olu iona y signi icance o genomic GC con en di e si y in
monoco s. P oc. Na l. Acad. Sci. U.S.A. , 111 , E4096–E4102.
38. Wang,Y., Mao,J.-M., Wang,G.-D., Luo,Z.-P., Yang,L., Yao,Q. and
Chen,K.-P. (2020) Human SARS-CoV-2 has e ol ed o educe CG
dinucleo ide in i s open eading ames. Sci. Rep. , 10 , 12331.
39. Ma y
´
a
ˇ
sek,R. and Ko a
ˇ
´
ık,A. (2020) Mu a ion pa e ns o human
SARS-CoV-2 and ba RaTG13 co ona i us genomes a e s ongly
biased owa ds C > U ansi ions, indica ing apid e olu ion in hei
Hos s. Genes (Basel) , 11 , 761.
40. Goswami,P., Ba as,M., Lexa,M., Boh
´
alo
´
a,N., Voln
´
a,A.,
ˇ
Ce e
ˇ
n,J.,
ˇ
Ce e
ˇ
no
´
a,V., Pe
ˇ
cinka,P.,
ˇ
Spunda,V., Foj a,M. e al. (2021)
SARS-CoV-2 ho -spo mu a ions a e signi ican ly en iched wi hin
in e ed epea s and CpG island loci. B ie Bioin o m , 22 , 1338–1345.
41. B
´
azda,V., Po ubiako
´
a,O., Can a a,A., Boh
´
alo
´
a,N., Cou al,J.,
Ba as,M., Foj a,M. and Me gny,J.-L. (2021) G-quad uplexes in
H1N1 in luenza genomes. BMC Genomics [Elec onic Resou ce] , 22 ,
77.
42. Boh
´
alo
´
a,N., Can a a,A., Ba as,M., Kau a,P.,
ˇ
S
ˇ
as n´y,J., Pe
ˇ
cinka,P.,
Foj a,M., Me gny,J.-L. and B
´
azda,V. (2021) Analyses o i al
genomes o G-quad uplex o ming sequences e eal hei co ela ion
wi h he ype o in ec ion. Biochimie , 186 , 13–27.
43. B own,J.C. (2007) High G+C con en o he pes simplex i us DNA:
p oposed ole in p o ec ion agains e o ansposon inse ion. Open
Biochem. J. , 1 , 33–42.
44. Vinog ado ,A.E. and Ana skaya,O.V. (2017) DNA helix: he
impo ance o being AT- ich. Mamm. Genome , 28 , 455–464.
45. Cal ignac-Spence ,S., D¨ux,A., Goga en,J.F. and Pa ono,L.V. (2021)
Chap e Two - Molecula a cheology o human i uses. In:
Kielian,M., Me enlei e ,T.C. and Roossinck,M.J. (eds.) Ad ances in
Vi us Resea ch . Academic P ess, Vol. 111 , pp. 31–61.
C
The Au ho (s) 2023. Published by Ox o d Uni e si y P ess on behal o Nucleic Acids Resea ch.
This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (h p: // c ea i ecommons.o g / licenses / by / 4.0 / ), which
pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium, p o ided he o iginal wo k is p ope ly ci ed.
Downloaded om h ps://academic.oup.com/na /a icle/51/14/7198/7217046 by Technical Uni e si y o B no use on 16 Feb ua y 2024