i uses
A icle
Iden i ica ion o a New HIV-1 BC In e sub ype Ci cula ing
Recombinan Fo m (CRF108_BC) in Spain
Ja ie E. Cañada 1, Elena Delgado 1, Ho acio Gil 1, Mónica Sánchez 1, Sonia Beni o 1, Elena Ga cía-Bodas 1,
Ca men Gómez-González 2, And és Canu -Blasco 2, Joseba Po u-Zapi ain 3, Es e Sáez de Adana 4,
Mi eia De la Peña 5, So ía Iba a 5, Gus a o Cilla 6, JoséAn onio I iba en 7, Ana Ma ínez-Sapiña 8
and Michael M. Thomson 1,*
Ci a ion: Cañada, J.E.; Delgado, E.;
Gil, H.; Sánchez, M.; Beni o, S.;
Ga cía-Bodas, E.; Gómez-González, C.;
Canu -Blasco, A.; Po u-Zapi ain, J.;
Sáez de Adana, E.; e al. Iden i ica ion
o a New HIV-1 BC In e sub ype
Ci cula ing Recombinan Fo m
(CRF108_BC) in Spain. Vi uses 2021,
13, 93. h ps://doi.o g/10.3390/
13010093
Academic Edi o s: William
M.M. Swi ze and
Dimi ios Pa aske is
Recei ed: 15 Decembe 2020
Accep ed: 8 Janua y 2021
Published: 12 Janua y 2021
Publishe ’s No e: MDPI s ays neu-
al wi h ega d o ju isdic ional clai-
ms in published maps and ins i u io-
nal a ilia ions.
Copy igh : © 2021 by he au ho s. Li-
censee MDPI, Basel, Swi ze land.
This a icle is an open access a icle
dis ibu ed unde he e ms and con-
di ions o he C ea i e Commons A -
ibu ion (CC BY) license (h ps://
c ea i ecommons.o g/licenses/by/
4.0/).
1HIV Biology and Va iabili y Uni , Cen o Nacional de Mic obiología, Ins i u o de Salud Ca los III,
Majadahonda, 28220 Mad id, Spain; [email p o ec ed] (J.E.C.); [email p o ec ed] (E.D.); [email p o ec ed] (H.G.);
[email p o ec ed] (M.S.); [email p o ec ed] (S.B.); [email p o ec ed] (E.G.-B.)
2Depa men o Mic obiology, Hospi al Uni e si a io A aba, 01009 Vi o ia-Gas eiz, Spain;
[email p o ec ed] (C.G.-G.); [email p o ec ed] (A.C.-B.)
3Bioa aba, In ec ious Diseases Resea ch G oup, 01009 Vi o ia-Gas eiz, Spain;
[email p o ec ed]
4Depa men o In ec ious Diseases-In e nal Medicine, Hospi al Uni e si a io A aba,
01009 Vi o ia-Gas eiz, Spain; es e [email p o ec ed]
5Depa men o In ec ious Diseases, Hospi al Uni e si a io Basu o, 48013 Bilbao, Spain;
mi eia.delapena igue [email p o ec ed] (M.D.l.P.); [email p o ec ed] (S.I.)
6Biodonos ia, Depa men o Mic obiology, Hospi al Uni e si a io Donos ia, 20080 San Sebas ián, Spain;
[email p o ec ed]
7
Biodonos ia, Depa men o In ec ious Diseases, Hospi al Uni e si a io Donos ia, 20080 San Sebas ián, Spain;
[email p o ec ed]
8Depa men o Mic obiology, Hospi al Uni e si a io Miguel Se e , 50009 Za agoza, Spain;
[email p o ec ed]
*Co espondence: [email p o ec ed]; Tel.: +34-918-223-900
Abs ac :
The ex ao dina y gene ic a iabili y o human immunode iciency i us ype 1 (HIV-1)
g oup M has led o he iden i ica ion o 10 sub ypes, 102 ci cula ing ecombinan o ms (CRFs) and
nume ous unique ecombinan o ms. Among CRFs, 11 de i ed om sub ypes B and C ha e been
iden i ied in China, B azil, and I aly. He e we iden i y a new HIV-1 CRF_BC in No he n Spain.
O iginally, a phylogene ic clus e o 15 i uses o sub ype C in p o ease- e e se ansc ip ase was
iden i ied in an HIV-1 molecula su eillance s udy in Spain, mos o hem om indi iduals om he
Basque Coun y and he e osexually ansmi ed. Analyses o nea ull-leng h genome sequences om
six i uses om h ee ci ies e ealed ha hey we e BC ecombinan wi h coinciden mosaic s uc u es
di e en om known CRFs. This allowed he de ini ion o a new HIV-1 CRF designa ed CRF108_BC,
whose genome is p edominan ly o sub ype C, wi h ou sho sub ype B agmen s. Phylogene ic
analyses wi h da abase sequences suppo ed a B azilian ances y o he pa en al sub ype C s ain.
Coalescen Bayesian analyses es ima ed he mos ecen common ances o o CRF108_BC in he ci y
o Vi o ia, Basque Coun y, a ound 2000. CRF108_BC is he i s CRF_BC iden i ied in Spain and he
second in Eu ope, a e CRF60_BC, bo h phylogene ically ela ed o B azilian sub ype C s ains.
Keywo ds:
HIV-1; ci cula ing ecombinan o m; HIV-1 gene ic di e si y; HIV-1 phylogeny; HIV-1
molecula epidemiology
1. In oduc ion
HIV-1 is cha ac e ized by high mu a ion and ecombina ion a es, which ha e led o
he gene a ion o ex ao dina y gene ic di e si y. Fou HIV-1 g oups ha e been cha ac-
e ized: M, N, O, and P. G oup M, he oldes lineage [
1
,
2
], is he causa i e o he global
pandemic and is subdi ided in o en sub ypes (A–D, F–H, J–L), o which sub ype C is
he mos p e alen wo ldwide, ci cula ing mainly in Sou he n and Eas A ica, India,
Vi uses 2021,13, 93. h ps://doi.o g/10.3390/ 13010093 h ps://www.mdpi.com/jou nal/ i uses
Vi uses 2021,13, 93 2 o 13
and Sou he n B azil, and sub ype B a e he mos p e alen in Wes e n Eu ope and he
Ame icas [3].
Recombina ion be ween sub ypes has led o he gene a ion o ci cula ing and unique
ecombinan o ms (CRFs and URFs, espec i ely), which ep esen a ound 23% o HIV-1
s ains wo ldwide [
3
]. To de ine a CRF, a leas h ee HIV-1 nea ull-leng h genomes
(NFLGs) mus be cha ac e ized om epidemiologically unlinked indi iduals, showing
iden ical mosaic pa e ns and clus e ing in phylogene ic ees apa om p e iously de ined
CRFs [4]. To da e, a o al o 102 CRFs ha e been epo ed in he li e a u e.
HIV-1 was in oduced in Wes e n Eu ope in he ea ly 1980s among men who ha e
sex wi h men (MSM) and pe sons who injec d ugs (PWID) in ec ed wi h i uses o
sub ype B, which is he cu en p edominan gene ic o m (67.2%), ollowed by sub ype C
(5.3%) [
5
]. Among non-sub ype B clades, he e a e se e al CRFs i s iden i ied in Wes e n
Eu ope: CRF04_cpx [
6
], CRF14_BG [
7
,
8
], CRF42_BF [
9
], CRF47_BF [
10
], CRF50_A1D [
11
],
CRF56_cpx [12], CRF60_BC [13,14], CRF73_BG [15], CRF94_cpx [16] and CRF98_06B [17].
In his s udy, we epo he i s CRF_BC iden i ied in Spain and he second in Eu ope,
es ima ing i s mos p obable o igin by phylogeog aphic and phylodynamic analyses.
2. Ma e ials and Me hods
2.1. Pa ien s
Plasma o whole blood samples we e collec ed in 1999–2020 om mo e han 13,000
HIV-1-in ec ed indi iduals om 10 Spanish egions o de e mina ion o an i e o i al
d ug esis ance mu a ions and o molecula epidemiological su eillance o HIV-1.
2.2. Nucleic Acid Ex ac ion, Ampli ica ion and Sequencing
RNA was ex ac ed om 1 mL plasma using NucliSENS
®
EasyMAG
®
(bioMé ieux,
Ma cy l’E oile, F ance). DNA was ex ac ed om 200
µ
L whole blood using QIAamp
®
DNA
DSP blood mini ki (Qiagen, Hilden, Ge many), ollowing he manu ac u e ’s ins uc ions.
A p o ease- e e se ansc ip ase (PR–RT) agmen o pol (HXB2 posi ions 2253–3629) was
ampli ied by RT–PCR ollowed by nes ed-PCR om RNA o by nes ed-PCR om DNA, as
p e iously desc ibed [18].
NFLGs we e ampli ied om plasma RNA, wi h PR–RT p e iously sequenced, h ough
RT–PCR ollowed by nes ed-PCR in i e o e lapping agmen s, as epo ed p e iously
in [
7
,
19
] and modi ied om [
20
]. The ampli ica ion s a egy is schema ically depic ed
in Figu e S1, and PCR p ime s a e lis ed in Table S1. Sequencing was done wi h an
au oma ed capilla y sequence . Fi een PR–RT, six NFLG, and one semigenome sequence
we e deposi ed in GenBank (Table 1).
Table 1. Epidemiological and clinical da a o he pa ien s and GenBank accessions o sequences.
Sample Ci y Region *
Coun y
o
O igin
Yea o
Diagno-
sis
Yea o
Collec-
ion
Gende Age T ansmission
Rou e †
PR–RT
GenBank
Accession
NFLG
GenBank
Accession
P1363
Vi o ia
Basque C.
Spain 2006 2010 M 64 He MT436238 -
P1607
Vi o ia
Basque C.
B azil 2007 2007 F 20 T ans MT436239 -
P2085
Vi o ia
Basque C.
Spain 2008 2008 M 40 He MT559129 -
P2782
San Sebas ián
Basque C.
Spain 2011 2011 M 38 He MT436240 -
P2969
San Sebas ián
Basque C.
Spain 2011 2011 F 32 He MT436241 -
P3536
Bilbao
Basque C.
Spain 2013 2013 M 70 He MT436242 MT559130
P4439
Vi o ia
Basque C.
Spain 2016 2016 F 58 He MT436244 MN172222
P4517
Vi o ia
Basque C.
Spain 2010 2016 M 58 n.a. MT436245 -
P4523
Bilbao
Basque C.
Spain 2016 2016 M 42 He MT436246 MT559131 ‡
P4697
Vi o ia
Basque C.
Spain 2017 2017 F 28 He MT436247 MT559132
P4719
Vi o ia
Basque C.
Spain 2017 2017 M 44 He MT436248 -
P4977
Vi o ia
Basque C.
Spain 2018 2018 M 53 He MT436249 MN172223
P5007
Vi o ia
Basque C.
Spain 2018 2018 M 35 He MT436250 MN172224
P5236
Vi o ia
Basque C.
Spain 2019 2019 F 65 He MT436251 -
Z0230
Za agoza A agon Spain 2017 2017 M 36 MSM MT436252 MN172225
* Basque.C.: Basque Coun y.
†
He : he e osexual; MSM: men who ha e sex wi h men; T ans: anssexual; n.a.: no a ailable.
‡
Semigenome.
Vi uses 2021,13, 93 3 o 13
2.3. Phylogene ic Analyses
Ini ial phylogene ic analyses we e pe o med wi h Fas T ee 2.1 [
21
] wi h mo e han
16,000 HIV-1 PR–RT sequences ob ained in ou labo a o y om mo e han 13,000 indi-
iduals whose samples we e collec ed in Spain in 1999–2020, simila sequences e ie ed
h ough BLAST sea ches [
22
] om he Los Alamos HIV Sequence Da abase [
23
], and
sub ype and CRF e e ences. Fo hese analyses, he gene al ime- e e sible wi h CAT
app oxima ion o a e he e ogenei y among si es (GTR + CAT) subs i u ion model was
used, wi h he assessmen o node suppo wi h Shimodai a–Hasegawa (SH)-like local
suppo alues.
Subsequen maximum-likelihood (ML) analyses we e pe o med wi h W-IQ-T ee [
24
],
including sequences simila o he iden i ied clus e , e ie ed om he HIV Sequence
Da abase [
23
] h ough BLAST sea ches [
22
]. These analyses we e pe o med wi h a 1200 n
agmen o he PR–RT egion (HXB2 posi ions 2253–3452). The subs i u ion model was
GTR wi h gamma-dis ibu ed he e ogenei y ac oss si es, allowing o a p opo ion o in-
a ian si es (GTR + G + I), and node suppo was assessed h ough ul a as boo s apping
wi h 1000 eplica es.
The ecombina ion pa e ns in NFLGs we e de e mined by boo scanning [
25
] us-
ing Simplo .3.5.1 [
26
], including HIV-1 sub ype e e ences downloaded om he Los
Alamos HIV Sequence Da abase, wi h a 250 n window mo ing in 20 n s eps and ees
cons uc ed wi h he neighbo -joining me hod and Kimu a 2-pa ame e subs i u ion model.
B eakpoin s we e loca ed mo e p ecisely h ough sequence inspec ion by de e mining
he segmen whe e simila i y o he BC ecombinan s wi h NFLG genomes o sub ype B
and B azilian sub ype C i uses changed be ween clades. B eakpoin s we e loca ed a
he midpoin be ween wo adjacen sub ype-disc imina ing nucleo ides (de ined as hose
di e ing be ween sub ype consensuses and p esen in >75% i uses o one o he pa en al
clades and in <10% o he o he ) whe e simila i y changed be ween sub ypes.
2.4. Phylogeog aphic and Phylodynamic Analyses
The ime and loca ion o he mos ecen common ances o (MRCA) we e es ima ed
wi h he Bayesian coalescen Ma ko Chain Mon e Ca lo (MCMC) me hod, implemen ed in
BEAST 1.10.4 [
27
], summa izing he se o ees o he pos e io dis ibu ion in a maximum
clade c edibili y (MCC) ee. PR–RT sequences (1.2 kb) o he iden i ied clus e we e used,
labeled wi h he loca ion and he yea o collec ion o he sample. Fi y sub ype C sequences
om di e en coun ies and collec ion yea s we e included o p o ide a empo al signal,
p e iously assessed wi h TempEs 1.5.1 [
28
]. We chose an HKY subs i u ion model wi h
gamma-dis ibu ed among-si e a e he e ogenei y and wo pa i ions in codon posi ions
(1s + 2nd; 3 d) [
29
]. A Bayesian skyline coalescen model was chosen wi h a logno mal
unco ela ed elaxed clock model. Uni o m p io s we e used o absolu e subs i u ion a es
(0–0.02 sub/si e/yea ). MCMC analyses we e un o 70 million gene a ions, sampling
e e y 4000 gene a ions. T ace 1.7.1 [
30
] was used o check MCMC con e gence, ensu ing
e ec i e sample sizes (ESS) o all pa ame e s abo e 200 [
30
]. T ees we e isualized wi h
FigT ee .1.3.1 (Rambau , h p:// ee.bio.ed.aC.uk/so wa e/ ig ee/).
2.5. An i e o i al D ug Resis ance Analysis
An i e o i al d ug esis ance was analyzed wi h he HIVdb p og am a S an o d
Uni e si y’s HIV D ug Resis ance Da abase [31].
3. Resul s
3.1. PR–RT Sequence Analyses
The phylogene ic analyses o HIV-1 PR–RT sequences om ou coho iden i ied a
monophyle ic clus e o 15 i uses o sub ype C suppo ed by an SH-like alue o 1, which
was designa ed C_2. The ML ee cons uc ed wi h W-IQ-T ee, including sequences e-
ie ed om da abases h ough BLAST simila i y sea ches, e ealed no addi ional i uses
b anching wi hin he clus e and i s ela ionship o i uses o he sub ype C s ain ci cu-
Vi uses 2021,13, 93 4 o 13
la ing in B azil (Figu e 1). Epidemiological and clinical da a o he 15 pa ien s o he C_2
clus e a e summa ized in Table 1. Mos o hem we e Spanish, excep a B azilian indi id-
ual, and we e diagnosed wi h HIV-1 in ec ion om 2006 o 2019 in he Basque Coun y
(10 in he ci y o Vi o ia), excep one pa ien diagnosed in Za agoza. Ten pa ien s we e men,
and 5 we e women; he e osexual ansmission was epo ed in 12 (80%) in ec ions, and one
pa ien was a sel - epo ed MSM.
Vi uses 2021, 13, x FOR PEER REVIEW 5 o 14
Figu e 1. Maximum likelihood ee o C_2 clus e . Simila sequences e ie ed om da abases and sub ype e e ences a e
also included. Only boo s ap alues ≥90% a e shown. Sequences ob ained in ou labo a o y a e in blue and bold ype.
Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he coun y o collec ion, i us names, and GenBank
accession. Sub ype e e ence sequences a e labeled wi h “Re .”.
No an i e o i al d ug esis ance mu a ions we e ound in any o he PR–RT se-
quences.
3.2. NFLG Sequence Analyses
To de e mine whe he he i uses g ouping in C_2 we e o uni o m sub ype along
hei genomes o ecombinan , six NFLG sequences and a semigenome ob ained om i-
uses collec ed in h ee ci ies we e analyzed by boo scanning. The analyses showed ha
he i uses we e BC ecombinan , wi h eigh b eakpoin s delimi ing ou sho sub ype B
agmen s, loca ed in pol, pol- i o e lap, i - p - a o e lap, and ne , espec i ely, in a ge-
nome p edominan ly o sub ype C (Figu e 2A). The mosaic s uc u e in e ed om he
Figu e 1.
Maximum likelihood ee o C_2 clus e . Simila sequences e ie ed om da abases and sub ype e e ences a e
also included. Only boo s ap alues
≥
90% a e shown. Sequences ob ained in ou labo a o y a e in blue and bold ype.
Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he coun y o collec ion, i us names, and GenBank
accession. Sub ype e e ence sequences a e labeled wi h “Re .”.
No an i e o i al d ug esis ance mu a ions we e ound in any o he PR–RT sequences.
Vi uses 2021,13, 93 5 o 13
3.2. NFLG Sequence Analyses
To de e mine whe he he i uses g ouping in C_2 we e o uni o m sub ype along
hei genomes o ecombinan , six NFLG sequences and a semigenome ob ained om
i uses collec ed in h ee ci ies we e analyzed by boo scanning. The analyses showed ha
he i uses we e BC ecombinan , wi h eigh b eakpoin s delimi ing ou sho sub ype
B agmen s, loca ed in pol,pol- i o e lap, i - p - a o e lap, and ne , espec i ely, in a
genome p edominan ly o sub ype C (Figu e 2A). The mosaic s uc u e in e ed om
he boo scan analyses complemen ed wi h sequence inspec ion o de ine mo e p ecisely
b eakpoin loca ions is shown in Figu e 2B. In an ML phylogene ic ee, all six newly
de i ed NFLG o he BC ecombinan i uses a e g ouped in a clus e sepa a e om
p e iously iden i ied CRF_BCs (Figu e 3). These esul s allow de ining a new HIV-1 CRF,
which was designa ed CRF108_BC. Compa ison o CRF108_BC’s mosaic s uc u e wi h
hose o p e iously iden i ied CRF_BCs is shown in Figu e S2.
Vi uses 2021, 13, x FOR PEER REVIEW 6 o 14
boo scan analyses complemen ed wi h sequence inspec ion o de ine mo e p ecisely
b eakpoin loca ions is shown in Figu e 2B. In an ML phylogene ic ee, all six newly de-
i ed NFLG o he BC ecombinan i uses a e g ouped in a clus e sepa a e om p e i-
ously iden i ied CRF_BCs (Figu e 3). These esul s allow de ining a new HIV-1 CRF,
which was designa ed CRF108_BC. Compa ison o CRF108_BC’s mosaic s uc u e wi h
hose o p e iously iden i ied CRF_BCs is shown in Figu e S2.
Figu e 2. (A) Boo scan analyses o 6 NFLG and 1 semigenome sequences o i uses o he C_2
clus e . The ho izon al axis ep esen s he posi ion in he HXB2 genome o he midpoin o a 250 n
window mo ing in 20 in inc emen s, and he e ical axis ep esen s he boo s ap alue suppo -
Figu e 2.
(
A
) Boo scan analyses o 6 NFLG and 1 semigenome sequences o i uses o he C_2 clus e . The ho izon al axis
ep esen s he posi ion in he HXB2 genome o he midpoin o a 250 n window mo ing in 20 in inc emen s, and he
Vi uses 2021,13, 93 6 o 13
e ical axis ep esen s he boo s ap alue suppo ing clus e ing o he que y sequence wi h sub ype e e ences. Ve ical dashed lines
deno e b eakpoin loca ions; (
B
) Mosaic s uc u e o HIV-1 BC in e sub ype ci cula ing ecombinan o m (CRF108_BC). B eakpoin
posi ions in he HXB2 genome a e indica ed.
Vi uses 2021, 13, x FOR PEER REVIEW 7 o 14
ing clus e ing o he que y sequence wi h sub ype e e ences. Ve ical dashed lines deno e b eak-
poin loca ions; (B) Mosaic s uc u e o HIV-1 BC in e sub ype ci cula ing ecombinan o m
(CRF108_BC). B eakpoin posi ions in he HXB2 genome a e indica ed.
Figu e 3. Maximum likelihood ee o NFLGs o CRF108_BC and all CRF_BCs iden i ied o da e. Only boo s ap alues
≥90% a e shown. CRF108_BC sequences a e in blue and bold ype. Da abase Scheme 108. BC, BLAST sea ches o simila
sequences we e done wi h sub ype C and B agmen s a he Los Alamos HIV Sequence Da abase. Wi h ega d o he
sub ype C agmen s, sea ches we e done wi h he wo la ges agmen s in gag-pol and a - e - pu-en -ne , espec i ely.
A phylogene ic ee wi h all sub ype C conca ena ed agmen s om he NFLGs o CRF108_BC, and mos simila da abase
sequences showed he closes ela ionship wi h he B azilian i us 02BR2022 om Sao Paulo (Figu e 4). Simila i y sea ches
wi h he ou sub ype B agmen s and subsequen phylogene ic analyses wi h indi idual o conca ena ed agmen s ailed
o iden i y any da abase i us ela ed o he sub ype B pa en al s ain o CRF108_BC.
Figu e 3.
Maximum likelihood ee o NFLGs o CRF108_BC and all CRF_BCs iden i ied o da e. Only boo s ap alues
≥
90% a e shown. CRF108_BC sequences a e in blue and bold ype. Da abase Scheme 108. BC, BLAST sea ches o simila
sequences we e done wi h sub ype C and B agmen s a he Los Alamos HIV Sequence Da abase. Wi h ega d o he
sub ype C agmen s, sea ches we e done wi h he wo la ges agmen s in gag-pol and a - e - pu-en -ne , espec i ely.
A phylogene ic ee wi h all sub ype C conca ena ed agmen s om he NFLGs o CRF108_BC, and mos simila da abase
sequences showed he closes ela ionship wi h he B azilian i us 02BR2022 om Sao Paulo (Figu e 4). Simila i y sea ches
wi h he ou sub ype B agmen s and subsequen phylogene ic analyses wi h indi idual o conca ena ed agmen s ailed
o iden i y any da abase i us ela ed o he sub ype B pa en al s ain o CRF108_BC.
Vi uses 2021,13, 93 7 o 13
Vi uses 2021, 13, x FOR PEER REVIEW 8 o 14
Figu e 4. Phylogene ic ee o conca ena ed sub ype C agmen s o CRF108_BC. Only boo s ap alues ≥90% a e shown.
CRF108_BC sequences a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he
coun y o sample collec ion, i us name, and GenBank accession.
3.3. Phylogeog aphic and Phylodynamic Analyses
To es ima e he empo al and geog aphic o igin o CRF108_BC, a Bayesian coalescen
analysis was pe o med wi h he 15 PR-RT sequences o he C_2 clus e and 50 sub ype C
da abase sequences om di e en coun ies o which yea and loca ion o sample collec-
ion we e a ailable. P io o his analysis, he exis ence o a empo al signal was checked
wi h TempEs 1.5.1, which e ealed a clock-like s uc u e in he da a se ( 2 = 0.38), indi-
ca ing su icien empo al signal o pe o m he analyses.
The MRCA o he clus e was es ima ed a ound 2000 (95% HPD, 1995–2004) in he
ci y o Vi o ia wi h a loca ion pos e io p obabili y o 0.998. An ances y in B azil was also
s ongly suppo ed (Figu e 5).
Figu e 4.
Phylogene ic ee o conca ena ed sub ype C agmen s o CRF108_BC. Only boo s ap alues
≥
90% a e shown.
CRF108_BC sequences a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he
coun y o sample collec ion, i us name, and GenBank accession.
3.3. Phylogeog aphic and Phylodynamic Analyses
To es ima e he empo al and geog aphic o igin o CRF108_BC, a Bayesian coalescen
analysis was pe o med wi h he 15 PR-RT sequences o he C_2 clus e and 50 sub ype
C da abase sequences om di e en coun ies o which yea and loca ion o sample
collec ion we e a ailable. P io o his analysis, he exis ence o a empo al signal was
checked wi h TempEs 1.5.1, which e ealed a clock-like s uc u e in he da a se (
2
= 0.38),
indica ing su icien empo al signal o pe o m he analyses.
The MRCA o he clus e was es ima ed a ound 2000 (95% HPD, 1995–2004) in he
ci y o Vi o ia wi h a loca ion pos e io p obabili y o 0.998. An ances y in B azil was also
s ongly suppo ed (Figu e 5).
Vi uses 2021,13, 93 8 o 13
Vi uses 2021, 13, x FOR PEER REVIEW 9 o 14
Figu e 5. Maximum clade c edibili y ee o PR–RT sequences o CRF108_BC and 50 sub ype C sequences om da abases.
Sequences belonging o he CRF108_BC clus e a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype,
wo-le e ISO coun y-code, i us name, and GenBank accession. Colo s o e minal and in e nal b anches ep esen sam-
pling loca ions and mos p obable loca ions o he co esponding nodes, espec i ely, acco ding o he legend. Fo he
nodes co esponding o he clus e and i s closes ances o , he mos p obable loca ions and he mean MRCA (wi h 95%
HPD in e als) a e indica ed. Nodes suppo ed by PP (pos e io p obabili y) = 0.998–1 and PP = 0.95–0.9979 a e ma ked
wi h illed and un illed ci cles, espec i ely. 29 b anches co esponding o sequences om Bo swana, Cyp us, E hiopia,
Geo gia, Is ael, Kenya, Malawi, Senegal, Sou h A ica, Sweden, Tanzania, Uni ed Kingdom, Yemen, and Zambia we e
collapsed o be e iewing.
4. Discussion
The cha ac e iza ion o six HIV-1 NFLG sequences o BC ecombinan i uses ob-
ained om epidemiologically-unlinked pa ien s showing a coinciden mosaic s uc u e
di e en om p e iously iden i ied CRFs and clus e ing wi h a 100% boo s ap alue al-
lowed o de ine a new CRF, designa ed CRF108_BC. Nine addi ional pa ial sequences a e
g ouped in a monophyle ic clus e in PR-RT. All 15 pa ien s we e diagnosed wi h HIV-1
in ec ion in No he n Spain be ween 2006 and 2019 and we e in ec ed p edominan ly ia
he e osexual con ac .
CRF108_BC is he i s CRF ecombinan o sub ypes B and C pa en al s ains iden i-
ied in Spain and he second in Eu ope (a e CRF60_BC in I aly [13]). Ten CRF_BCs ha e
been iden i ied elsewhe e, nine in China [CRF07_BC [32], CRF08_BC [33], CRF57_BC [34],
CRF61_BC [35], CRF62_BC [36], CRF64_BC [37], CRF85_BC [38], CRF86_BC [39] and
CRF88_BC [40]] and one in B azil [CRF31_BC [41]]. Simila o CRF60_BC, he sub ype C
pa en al s ain o CRF108_BC is phylogene ically ela ed o he sub ype C s ain ci cula -
ing in B azil. Howe e , he ecombina ion e en gi ing ise o CRF108_BC could ha e
Figu e 5.
Maximum clade c edibili y ee o PR–RT sequences o CRF108_BC and 50 sub ype C sequences om da abases.
Sequences belonging o he CRF108_BC clus e a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype,
wo-le e ISO coun y-code, i us name, and GenBank accession. Colo s o e minal and in e nal b anches ep esen
sampling loca ions and mos p obable loca ions o he co esponding nodes, espec i ely, acco ding o he legend. Fo he
nodes co esponding o he clus e and i s closes ances o , he mos p obable loca ions and he mean MRCA (wi h 95%
HPD in e als) a e indica ed. Nodes suppo ed by PP (pos e io p obabili y) = 0.998–1 and PP = 0.95–0.9979 a e ma ked
wi h illed and un illed ci cles, espec i ely. 29 b anches co esponding o sequences om Bo swana, Cyp us, E hiopia,
Geo gia, Is ael, Kenya, Malawi, Senegal, Sou h A ica, Sweden, Tanzania, Uni ed Kingdom, Yemen, and Zambia we e
collapsed o be e iewing.
4. Discussion
The cha ac e iza ion o six HIV-1 NFLG sequences o BC ecombinan i uses ob ained
om epidemiologically-unlinked pa ien s showing a coinciden mosaic s uc u e di e en
om p e iously iden i ied CRFs and clus e ing wi h a 100% boo s ap alue allowed o
de ine a new CRF, designa ed CRF108_BC. Nine addi ional pa ial sequences a e g ouped
in a monophyle ic clus e in PR-RT. All 15 pa ien s we e diagnosed wi h HIV-1 in ec ion in
No he n Spain be ween 2006 and 2019 and we e in ec ed p edominan ly ia he e osexual
con ac .
CRF108_BC is he i s CRF ecombinan o sub ypes B and C pa en al s ains iden i ied
in Spain and he second in Eu ope (a e CRF60_BC in I aly [
13
]). Ten CRF_BCs ha e been
iden i ied elsewhe e, nine in China [CRF07_BC [
32
], CRF08_BC [
33
], CRF57_BC [
34
],
CRF61_BC [
35
], CRF62_BC [
36
], CRF64_BC [
37
], CRF85_BC [
38
], CRF86_BC [
39
] and
CRF88_BC [
40
]] and one in B azil [CRF31_BC [
41
]]. Simila o CRF60_BC, he sub ype C
pa en al s ain o CRF108_BC is phylogene ically ela ed o he sub ype C s ain ci cula ing
in B azil. Howe e , he ecombina ion e en gi ing ise o CRF108_BC could ha e occu ed
Vi uses 2021,13, 93 9 o 13
ei he in B azil, whe e bo h B and C sub ypes co-ci cula e a high p opo ions in some
a eas [
42
,
43
] o in Spain, since we could no ack he ances y o he pa en al sub ype B
s ain o any coun y. A Spanish o igin o CRF108_BC would be suppo ed by an in e ed
Spanish MRCA in he ci y o Vi o ia, Basque Coun y, No he n Spain, a ound 2000.
Howe e , we canno ule ou ha CRF108_BC could be ci cula ing a a low p e alence in
some a ea(s) o B azil, which could explain i s lack o ep esen a ion in public sequence
da abases.
CRF108_BC, simila ly o all o he CRF_BCs iden i ied o da e, is p edominan ly o
sub ype C. I is in e es ing o no e ha in all CRF_BCs, en is mos ly o sub ype C, which
could lead o he specula ion ha a sub ype C en elope could con e e olu iona y ad an-
ages ega ding i al i ness, escape o immune esponse o a mo e e icien eplica ion o
ansmission o e a sub ype B en elope. Fu he in es iga ions a e equi ed o con i m
any o hese possibili ies.
To da e, he expansion o CRF108_BC has aken place om 2006 o 2019, wi h 53%
o cases diagnosed in he las 4 yea s, and limi ed o a small geog aphical a ea in Spain,
wi h 10 o 15 cases in he ci y o Vi o ia and 4 o he cases in wo neighbo ing p o inces o
he Basque Coun y. Ou side o he Basque Coun y, we only ha e ound a single in ec ion
wi h CRF108_BC in he ci y o Za agoza among HIV-1 sequences om mo e han 13,000
pa ien s om 10 egions o Spain analyzed by us and all sequences om public da abases.
In Asia, he majo i y o CRF_BCs we e ansmi ed ini ially among pe sons who injec
d ugs (PWID) [
40
,
44
] and ecen ly expanded among MSM [
45
]. In Eu ope, CRF60_BC
is associa ed wi h p opaga ion among MSM [
13
]. Al hough, a p esen , MSM is he
mos equen ansmission ou e o HIV-1 in Spain [
46
], wo CRFs p e iously iden i ied
in Spain, CRF14_BG and CRF73_BG, p opaga ed mainly among PWID [
7
,
15
]. A hi d
CRF iden i ied in Spain, CRF47_BF, was associa ed wi h he e osexual ansmission [
10
],
simila ly o CRF108_BC. Howe e , we ha e obse ed u he p opaga ion o CRF47_BF
among MSM [
47
]. This pa e n could po en ially be epea ed wi h CRF108_BC since, along
wi h he e osexually ansmi ed cases (wi h 27% women), we ind an MSM diagnosed in
2017. This sugges s ha ansmission ne wo ks o HIV-1 in Spain among he e osexuals
may be d i ing owa ds MSM. We ha e also obse ed he e e se si ua ion in he case o a
CRF02_AG clus e sp eading om MSM o a he e osexual ne wo k [48].
NFLG sequencing o HIV-1 s ains in ol ed in expanding ansmission clus e s is
highly ecommendable in o de o iden i y new CRFs, which may ha e acqui ed adap i e
ad an ages h ough ecombina ion [
49
,
50
] e en when ecombina ion is no suspec ed
in pa ial sequences, as i occu s wi h CRF108_BC. Gene ic di e si y and ecombina ion
a e majo obs acles o he de elopmen o an e ec i e accine agains HIV-1 [
51
–
53
].
Cha ac e iza ion and molecula epidemiological su eillance o expanding new HIV-1
CRFs may play an impo an ole in public heal h ac ions, including he selec ion o op imal
immunogens o e ec i e accines.
5. Conclusions
A new HIV-1 CRF de i ed om sub ypes B and C, designa ed CRF108_BC, has been
iden i ied a e he analysis o six NFLG om h ee ci ies in No he n Spain, which was
o iginally iden i ied as a sub ype C clus e in PR–RT sequences comp ising 15 indi iduals.
Phylogene ic and phylogeog aphic analyses poin o a B azilian ances y, al hough i is
unclea whe he he ecombina ion e en ook place in B azil o in Spain.
Among CRFs de i ed om B and C sub ypes, CRF108_BC is he 12 h iden i ied,
he i s in Spain and he second in Eu ope. I s sp ead is cu en ly limi ed, wi h only
15 cases de ec ed, mos o hem ansmi ed ia he e osexual con ac . Conside ing ha
mo e han hal o hem we e diagnosed in he las ou yea s, molecula epidemiological
su eillance seems jus i ied o examine u he sp ead. The esul s o his s udy also
ad oca e o NFLG sequence cha ac e iza ion o eme ging HIV-1 clus e s, which may
ep esen new CRFs.