scieee Open visual document viewer

Identification of a New HIV-1 BC Intersubtype Circulating Recombinant Form (CRF108_BC) in Spain

Cañada, J.E.; Thomson, M.M.; Iribarren, J.A.; Benito, S.; Sáez de Adana, E.; Gil, H.; Sánchez, M.; Cilla, G.; Delgado, E.; Gómez-González, C.; Portu-Zapirain, J.; García-Bodas, E.; Martínez-Sapiña, A.; De la Peña, M.; Ibarra, S.; Canut-Blasco, A.

Abstract

The extraordinary genetic variability of human immunodeficiency virus type 1 (HIV-1) group M has led to the identification of 10 subtypes, 102 circulating recombinant forms (CRFs) and numerous unique recombinant forms. Among CRFs, 11 derived from subtypes B and C have been identified in China, Brazil, and Italy. Here we identify a new HIV-1 CRF_BC in Northern Spain. Originally, a phylogenetic cluster of 15 viruses of subtype C in protease-reverse transcriptase was identified in an HIV-1 molecular surveillance study in Spain, most of them from individuals from the Basque Country and heterosexually transmitted. Analyses of near full-length genome sequences from six viruses from three cities revealed that they were BC recombinant with coincident mosaic structures different from known CRFs. This allowed the definition of a new HIV-1 CRF designated CRF108_BC, whose genome is predominantly of subtype C, with four short subtype B fragments. Phylogenetic analyses with database sequences supported a Brazilian ancestry of the parental subtype C strain. Coalescent Bayesian analyses estimated the most recent common ancestor of CRF108_BC in the city of Vitoria, Basque Country, around 2000. CRF108_BC is the first CRF_BC identified in Spain and the second in Europe, after CRF60_BC, both phylogenetically related to Brazilian subtype C strains. Cañada, J.E.; Delgado, E.; Gil, H.; Sánchez, M.; Benito, S.; García-Bodas, E.; Gómez-González, C.; Canut-Blasco, A.; Portu-Zapirain, J.; Sáez de Adana, E.; De la Peña, M.; Ibarra, S.; Cilla, G.; Iribarren, J.A.; Martínez-Sapiña, A.; Thomson, M.M.

Full text

i uses A icle Iden i ica ion o a New HIV-1 BC In e sub ype Ci cula ing Recombinan Fo m (CRF108_BC) in Spain Ja ie E. Cañada 1, Elena Delgado 1, Ho acio Gil 1, Mónica Sánchez 1, Sonia Beni o 1, Elena Ga cía-Bodas 1, Ca men Gómez-González 2, And és Canu -Blasco 2, Joseba Po u-Zapi ain 3, Es e Sáez de Adana 4, Mi eia De la Peña 5, So ía Iba a 5, Gus a o Cilla 6, JoséAn onio I iba en 7, Ana Ma ínez-Sapiña 8 and Michael M. Thomson 1,*   Ci a ion: Cañada, J.E.; Delgado, E.; Gil, H.; Sánchez, M.; Beni o, S.; Ga cía-Bodas, E.; Gómez-González, C.; Canu -Blasco, A.; Po u-Zapi ain, J.; Sáez de Adana, E.; e al. Iden i ica ion o a New HIV-1 BC In e sub ype Ci cula ing Recombinan Fo m (CRF108_BC) in Spain. Vi uses 2021, 13, 93. h ps://doi.o g/10.3390/ 13010093 Academic Edi o s: William M.M. Swi ze and Dimi ios Pa aske is Recei ed: 15 Decembe 2020 Accep ed: 8 Janua y 2021 Published: 12 Janua y 2021 Publishe ’s No e: MDPI s ays neu- al wi h ega d o ju isdic ional clai- ms in published maps and ins i u io- nal a ilia ions. Copy igh : © 2021 by he au ho s. Li- censee MDPI, Basel, Swi ze land. This a icle is an open access a icle dis ibu ed unde he e ms and con- di ions o he C ea i e Commons A - ibu ion (CC BY) license (h ps:// c ea i ecommons.o g/licenses/by/ 4.0/). 1HIV Biology and Va iabili y Uni , Cen o Nacional de Mic obiología, Ins i u o de Salud Ca los III, Majadahonda, 28220 Mad id, Spain; [email p o ec ed] (J.E.C.); [email p o ec ed] (E.D.); [email p o ec ed] (H.G.); [email p o ec ed] (M.S.); [email p o ec ed] (S.B.); [email p o ec ed] (E.G.-B.) 2Depa men o Mic obiology, Hospi al Uni e si a io A aba, 01009 Vi o ia-Gas eiz, Spain; [email p o ec ed] (C.G.-G.); [email p o ec ed] (A.C.-B.) 3Bioa aba, In ec ious Diseases Resea ch G oup, 01009 Vi o ia-Gas eiz, Spain; [email p o ec ed] 4Depa men o In ec ious Diseases-In e nal Medicine, Hospi al Uni e si a io A aba, 01009 Vi o ia-Gas eiz, Spain; es e [email p o ec ed] 5Depa men o In ec ious Diseases, Hospi al Uni e si a io Basu o, 48013 Bilbao, Spain; mi eia.delapena igue [email p o ec ed] (M.D.l.P.); [email p o ec ed] (S.I.) 6Biodonos ia, Depa men o Mic obiology, Hospi al Uni e si a io Donos ia, 20080 San Sebas ián, Spain; [email p o ec ed] 7 Biodonos ia, Depa men o In ec ious Diseases, Hospi al Uni e si a io Donos ia, 20080 San Sebas ián, Spain; [email p o ec ed] 8Depa men o Mic obiology, Hospi al Uni e si a io Miguel Se e , 50009 Za agoza, Spain; [email p o ec ed] *Co espondence: [email p o ec ed]; Tel.: +34-918-223-900 Abs ac : The ex ao dina y gene ic a iabili y o human immunode iciency i us ype 1 (HIV-1) g oup M has led o he iden i ica ion o 10 sub ypes, 102 ci cula ing ecombinan o ms (CRFs) and nume ous unique ecombinan o ms. Among CRFs, 11 de i ed om sub ypes B and C ha e been iden i ied in China, B azil, and I aly. He e we iden i y a new HIV-1 CRF_BC in No he n Spain. O iginally, a phylogene ic clus e o 15 i uses o sub ype C in p o ease- e e se ansc ip ase was iden i ied in an HIV-1 molecula su eillance s udy in Spain, mos o hem om indi iduals om he Basque Coun y and he e osexually ansmi ed. Analyses o nea ull-leng h genome sequences om six i uses om h ee ci ies e ealed ha hey we e BC ecombinan wi h coinciden mosaic s uc u es di e en om known CRFs. This allowed he de ini ion o a new HIV-1 CRF designa ed CRF108_BC, whose genome is p edominan ly o sub ype C, wi h ou sho sub ype B agmen s. Phylogene ic analyses wi h da abase sequences suppo ed a B azilian ances y o he pa en al sub ype C s ain. Coalescen Bayesian analyses es ima ed he mos ecen common ances o o CRF108_BC in he ci y o Vi o ia, Basque Coun y, a ound 2000. CRF108_BC is he i s CRF_BC iden i ied in Spain and he second in Eu ope, a e CRF60_BC, bo h phylogene ically ela ed o B azilian sub ype C s ains. Keywo ds: HIV-1; ci cula ing ecombinan o m; HIV-1 gene ic di e si y; HIV-1 phylogeny; HIV-1 molecula epidemiology 1. In oduc ion HIV-1 is cha ac e ized by high mu a ion and ecombina ion a es, which ha e led o he gene a ion o ex ao dina y gene ic di e si y. Fou HIV-1 g oups ha e been cha ac- e ized: M, N, O, and P. G oup M, he oldes lineage [ 1 , 2 ], is he causa i e o he global pandemic and is subdi ided in o en sub ypes (A–D, F–H, J–L), o which sub ype C is he mos p e alen wo ldwide, ci cula ing mainly in Sou he n and Eas A ica, India, Vi uses 2021,13, 93. h ps://doi.o g/10.3390/ 13010093 h ps://www.mdpi.com/jou nal/ i uses Vi uses 2021,13, 93 2 o 13 and Sou he n B azil, and sub ype B a e he mos p e alen in Wes e n Eu ope and he Ame icas [3]. Recombina ion be ween sub ypes has led o he gene a ion o ci cula ing and unique ecombinan o ms (CRFs and URFs, espec i ely), which ep esen a ound 23% o HIV-1 s ains wo ldwide [ 3 ]. To de ine a CRF, a leas h ee HIV-1 nea ull-leng h genomes (NFLGs) mus be cha ac e ized om epidemiologically unlinked indi iduals, showing iden ical mosaic pa e ns and clus e ing in phylogene ic ees apa om p e iously de ined CRFs [4]. To da e, a o al o 102 CRFs ha e been epo ed in he li e a u e. HIV-1 was in oduced in Wes e n Eu ope in he ea ly 1980s among men who ha e sex wi h men (MSM) and pe sons who injec d ugs (PWID) in ec ed wi h i uses o sub ype B, which is he cu en p edominan gene ic o m (67.2%), ollowed by sub ype C (5.3%) [ 5 ]. Among non-sub ype B clades, he e a e se e al CRFs i s iden i ied in Wes e n Eu ope: CRF04_cpx [ 6 ], CRF14_BG [ 7 , 8 ], CRF42_BF [ 9 ], CRF47_BF [ 10 ], CRF50_A1D [ 11 ], CRF56_cpx [12], CRF60_BC [13,14], CRF73_BG [15], CRF94_cpx [16] and CRF98_06B [17]. In his s udy, we epo he i s CRF_BC iden i ied in Spain and he second in Eu ope, es ima ing i s mos p obable o igin by phylogeog aphic and phylodynamic analyses. 2. Ma e ials and Me hods 2.1. Pa ien s Plasma o whole blood samples we e collec ed in 1999–2020 om mo e han 13,000 HIV-1-in ec ed indi iduals om 10 Spanish egions o de e mina ion o an i e o i al d ug esis ance mu a ions and o molecula epidemiological su eillance o HIV-1. 2.2. Nucleic Acid Ex ac ion, Ampli ica ion and Sequencing RNA was ex ac ed om 1 mL plasma using NucliSENS ® EasyMAG ® (bioMé ieux, Ma cy l’E oile, F ance). DNA was ex ac ed om 200 µ L whole blood using QIAamp ® DNA DSP blood mini ki (Qiagen, Hilden, Ge many), ollowing he manu ac u e ’s ins uc ions. A p o ease- e e se ansc ip ase (PR–RT) agmen o pol (HXB2 posi ions 2253–3629) was ampli ied by RT–PCR ollowed by nes ed-PCR om RNA o by nes ed-PCR om DNA, as p e iously desc ibed [18]. NFLGs we e ampli ied om plasma RNA, wi h PR–RT p e iously sequenced, h ough RT–PCR ollowed by nes ed-PCR in i e o e lapping agmen s, as epo ed p e iously in [ 7 , 19 ] and modi ied om [ 20 ]. The ampli ica ion s a egy is schema ically depic ed in Figu e S1, and PCR p ime s a e lis ed in Table S1. Sequencing was done wi h an au oma ed capilla y sequence . Fi een PR–RT, six NFLG, and one semigenome sequence we e deposi ed in GenBank (Table 1). Table 1. Epidemiological and clinical da a o he pa ien s and GenBank accessions o sequences. Sample Ci y Region * Coun y o O igin Yea o Diagno- sis Yea o Collec- ion Gende Age T ansmission Rou e † PR–RT GenBank Accession NFLG GenBank Accession P1363 Vi o ia Basque C. Spain 2006 2010 M 64 He MT436238 - P1607 Vi o ia Basque C. B azil 2007 2007 F 20 T ans MT436239 - P2085 Vi o ia Basque C. Spain 2008 2008 M 40 He MT559129 - P2782 San Sebas ián Basque C. Spain 2011 2011 M 38 He MT436240 - P2969 San Sebas ián Basque C. Spain 2011 2011 F 32 He MT436241 - P3536 Bilbao Basque C. Spain 2013 2013 M 70 He MT436242 MT559130 P4439 Vi o ia Basque C. Spain 2016 2016 F 58 He MT436244 MN172222 P4517 Vi o ia Basque C. Spain 2010 2016 M 58 n.a. MT436245 - P4523 Bilbao Basque C. Spain 2016 2016 M 42 He MT436246 MT559131 ‡ P4697 Vi o ia Basque C. Spain 2017 2017 F 28 He MT436247 MT559132 P4719 Vi o ia Basque C. Spain 2017 2017 M 44 He MT436248 - P4977 Vi o ia Basque C. Spain 2018 2018 M 53 He MT436249 MN172223 P5007 Vi o ia Basque C. Spain 2018 2018 M 35 He MT436250 MN172224 P5236 Vi o ia Basque C. Spain 2019 2019 F 65 He MT436251 - Z0230 Za agoza A agon Spain 2017 2017 M 36 MSM MT436252 MN172225 * Basque.C.: Basque Coun y. † He : he e osexual; MSM: men who ha e sex wi h men; T ans: anssexual; n.a.: no a ailable. ‡ Semigenome. Vi uses 2021,13, 93 3 o 13 2.3. Phylogene ic Analyses Ini ial phylogene ic analyses we e pe o med wi h Fas T ee 2.1 [ 21 ] wi h mo e han 16,000 HIV-1 PR–RT sequences ob ained in ou labo a o y om mo e han 13,000 indi- iduals whose samples we e collec ed in Spain in 1999–2020, simila sequences e ie ed h ough BLAST sea ches [ 22 ] om he Los Alamos HIV Sequence Da abase [ 23 ], and sub ype and CRF e e ences. Fo hese analyses, he gene al ime- e e sible wi h CAT app oxima ion o a e he e ogenei y among si es (GTR + CAT) subs i u ion model was used, wi h he assessmen o node suppo wi h Shimodai a–Hasegawa (SH)-like local suppo alues. Subsequen maximum-likelihood (ML) analyses we e pe o med wi h W-IQ-T ee [ 24 ], including sequences simila o he iden i ied clus e , e ie ed om he HIV Sequence Da abase [ 23 ] h ough BLAST sea ches [ 22 ]. These analyses we e pe o med wi h a 1200 n agmen o he PR–RT egion (HXB2 posi ions 2253–3452). The subs i u ion model was GTR wi h gamma-dis ibu ed he e ogenei y ac oss si es, allowing o a p opo ion o in- a ian si es (GTR + G + I), and node suppo was assessed h ough ul a as boo s apping wi h 1000 eplica es. The ecombina ion pa e ns in NFLGs we e de e mined by boo scanning [ 25 ] us- ing Simplo .3.5.1 [ 26 ], including HIV-1 sub ype e e ences downloaded om he Los Alamos HIV Sequence Da abase, wi h a 250 n window mo ing in 20 n s eps and ees cons uc ed wi h he neighbo -joining me hod and Kimu a 2-pa ame e subs i u ion model. B eakpoin s we e loca ed mo e p ecisely h ough sequence inspec ion by de e mining he segmen whe e simila i y o he BC ecombinan s wi h NFLG genomes o sub ype B and B azilian sub ype C i uses changed be ween clades. B eakpoin s we e loca ed a he midpoin be ween wo adjacen sub ype-disc imina ing nucleo ides (de ined as hose di e ing be ween sub ype consensuses and p esen in >75% i uses o one o he pa en al clades and in <10% o he o he ) whe e simila i y changed be ween sub ypes. 2.4. Phylogeog aphic and Phylodynamic Analyses The ime and loca ion o he mos ecen common ances o (MRCA) we e es ima ed wi h he Bayesian coalescen Ma ko Chain Mon e Ca lo (MCMC) me hod, implemen ed in BEAST 1.10.4 [ 27 ], summa izing he se o ees o he pos e io dis ibu ion in a maximum clade c edibili y (MCC) ee. PR–RT sequences (1.2 kb) o he iden i ied clus e we e used, labeled wi h he loca ion and he yea o collec ion o he sample. Fi y sub ype C sequences om di e en coun ies and collec ion yea s we e included o p o ide a empo al signal, p e iously assessed wi h TempEs 1.5.1 [ 28 ]. We chose an HKY subs i u ion model wi h gamma-dis ibu ed among-si e a e he e ogenei y and wo pa i ions in codon posi ions (1s + 2nd; 3 d) [ 29 ]. A Bayesian skyline coalescen model was chosen wi h a logno mal unco ela ed elaxed clock model. Uni o m p io s we e used o absolu e subs i u ion a es (0–0.02 sub/si e/yea ). MCMC analyses we e un o 70 million gene a ions, sampling e e y 4000 gene a ions. T ace 1.7.1 [ 30 ] was used o check MCMC con e gence, ensu ing e ec i e sample sizes (ESS) o all pa ame e s abo e 200 [ 30 ]. T ees we e isualized wi h FigT ee .1.3.1 (Rambau , h p:// ee.bio.ed.aC.uk/so wa e/ ig ee/). 2.5. An i e o i al D ug Resis ance Analysis An i e o i al d ug esis ance was analyzed wi h he HIVdb p og am a S an o d Uni e si y’s HIV D ug Resis ance Da abase [31]. 3. Resul s 3.1. PR–RT Sequence Analyses The phylogene ic analyses o HIV-1 PR–RT sequences om ou coho iden i ied a monophyle ic clus e o 15 i uses o sub ype C suppo ed by an SH-like alue o 1, which was designa ed C_2. The ML ee cons uc ed wi h W-IQ-T ee, including sequences e- ie ed om da abases h ough BLAST simila i y sea ches, e ealed no addi ional i uses b anching wi hin he clus e and i s ela ionship o i uses o he sub ype C s ain ci cu- Vi uses 2021,13, 93 4 o 13 la ing in B azil (Figu e 1). Epidemiological and clinical da a o he 15 pa ien s o he C_2 clus e a e summa ized in Table 1. Mos o hem we e Spanish, excep a B azilian indi id- ual, and we e diagnosed wi h HIV-1 in ec ion om 2006 o 2019 in he Basque Coun y (10 in he ci y o Vi o ia), excep one pa ien diagnosed in Za agoza. Ten pa ien s we e men, and 5 we e women; he e osexual ansmission was epo ed in 12 (80%) in ec ions, and one pa ien was a sel - epo ed MSM. Vi uses 2021, 13, x FOR PEER REVIEW 5 o 14 Figu e 1. Maximum likelihood ee o C_2 clus e . Simila sequences e ie ed om da abases and sub ype e e ences a e also included. Only boo s ap alues ≥90% a e shown. Sequences ob ained in ou labo a o y a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he coun y o collec ion, i us names, and GenBank accession. Sub ype e e ence sequences a e labeled wi h “Re .”. No an i e o i al d ug esis ance mu a ions we e ound in any o he PR–RT se- quences. 3.2. NFLG Sequence Analyses To de e mine whe he he i uses g ouping in C_2 we e o uni o m sub ype along hei genomes o ecombinan , six NFLG sequences and a semigenome ob ained om i- uses collec ed in h ee ci ies we e analyzed by boo scanning. The analyses showed ha he i uses we e BC ecombinan , wi h eigh b eakpoin s delimi ing ou sho sub ype B agmen s, loca ed in pol, pol- i o e lap, i - p - a o e lap, and ne , espec i ely, in a ge- nome p edominan ly o sub ype C (Figu e 2A). The mosaic s uc u e in e ed om he Figu e 1. Maximum likelihood ee o C_2 clus e . Simila sequences e ie ed om da abases and sub ype e e ences a e also included. Only boo s ap alues ≥ 90% a e shown. Sequences ob ained in ou labo a o y a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he coun y o collec ion, i us names, and GenBank accession. Sub ype e e ence sequences a e labeled wi h “Re .”. No an i e o i al d ug esis ance mu a ions we e ound in any o he PR–RT sequences. Vi uses 2021,13, 93 5 o 13 3.2. NFLG Sequence Analyses To de e mine whe he he i uses g ouping in C_2 we e o uni o m sub ype along hei genomes o ecombinan , six NFLG sequences and a semigenome ob ained om i uses collec ed in h ee ci ies we e analyzed by boo scanning. The analyses showed ha he i uses we e BC ecombinan , wi h eigh b eakpoin s delimi ing ou sho sub ype B agmen s, loca ed in pol,pol- i o e lap, i - p - a o e lap, and ne , espec i ely, in a genome p edominan ly o sub ype C (Figu e 2A). The mosaic s uc u e in e ed om he boo scan analyses complemen ed wi h sequence inspec ion o de ine mo e p ecisely b eakpoin loca ions is shown in Figu e 2B. In an ML phylogene ic ee, all six newly de i ed NFLG o he BC ecombinan i uses a e g ouped in a clus e sepa a e om p e iously iden i ied CRF_BCs (Figu e 3). These esul s allow de ining a new HIV-1 CRF, which was designa ed CRF108_BC. Compa ison o CRF108_BC’s mosaic s uc u e wi h hose o p e iously iden i ied CRF_BCs is shown in Figu e S2. Vi uses 2021, 13, x FOR PEER REVIEW 6 o 14 boo scan analyses complemen ed wi h sequence inspec ion o de ine mo e p ecisely b eakpoin loca ions is shown in Figu e 2B. In an ML phylogene ic ee, all six newly de- i ed NFLG o he BC ecombinan i uses a e g ouped in a clus e sepa a e om p e i- ously iden i ied CRF_BCs (Figu e 3). These esul s allow de ining a new HIV-1 CRF, which was designa ed CRF108_BC. Compa ison o CRF108_BC’s mosaic s uc u e wi h hose o p e iously iden i ied CRF_BCs is shown in Figu e S2. Figu e 2. (A) Boo scan analyses o 6 NFLG and 1 semigenome sequences o i uses o he C_2 clus e . The ho izon al axis ep esen s he posi ion in he HXB2 genome o he midpoin o a 250 n window mo ing in 20 in inc emen s, and he e ical axis ep esen s he boo s ap alue suppo - Figu e 2. ( A ) Boo scan analyses o 6 NFLG and 1 semigenome sequences o i uses o he C_2 clus e . The ho izon al axis ep esen s he posi ion in he HXB2 genome o he midpoin o a 250 n window mo ing in 20 in inc emen s, and he Vi uses 2021,13, 93 6 o 13 e ical axis ep esen s he boo s ap alue suppo ing clus e ing o he que y sequence wi h sub ype e e ences. Ve ical dashed lines deno e b eakpoin loca ions; ( B ) Mosaic s uc u e o HIV-1 BC in e sub ype ci cula ing ecombinan o m (CRF108_BC). B eakpoin posi ions in he HXB2 genome a e indica ed. Vi uses 2021, 13, x FOR PEER REVIEW 7 o 14 ing clus e ing o he que y sequence wi h sub ype e e ences. Ve ical dashed lines deno e b eak- poin loca ions; (B) Mosaic s uc u e o HIV-1 BC in e sub ype ci cula ing ecombinan o m (CRF108_BC). B eakpoin posi ions in he HXB2 genome a e indica ed. Figu e 3. Maximum likelihood ee o NFLGs o CRF108_BC and all CRF_BCs iden i ied o da e. Only boo s ap alues ≥90% a e shown. CRF108_BC sequences a e in blue and bold ype. Da abase Scheme 108. BC, BLAST sea ches o simila sequences we e done wi h sub ype C and B agmen s a he Los Alamos HIV Sequence Da abase. Wi h ega d o he sub ype C agmen s, sea ches we e done wi h he wo la ges agmen s in gag-pol and a - e - pu-en -ne , espec i ely. A phylogene ic ee wi h all sub ype C conca ena ed agmen s om he NFLGs o CRF108_BC, and mos simila da abase sequences showed he closes ela ionship wi h he B azilian i us 02BR2022 om Sao Paulo (Figu e 4). Simila i y sea ches wi h he ou sub ype B agmen s and subsequen phylogene ic analyses wi h indi idual o conca ena ed agmen s ailed o iden i y any da abase i us ela ed o he sub ype B pa en al s ain o CRF108_BC. Figu e 3. Maximum likelihood ee o NFLGs o CRF108_BC and all CRF_BCs iden i ied o da e. Only boo s ap alues ≥ 90% a e shown. CRF108_BC sequences a e in blue and bold ype. Da abase Scheme 108. BC, BLAST sea ches o simila sequences we e done wi h sub ype C and B agmen s a he Los Alamos HIV Sequence Da abase. Wi h ega d o he sub ype C agmen s, sea ches we e done wi h he wo la ges agmen s in gag-pol and a - e - pu-en -ne , espec i ely. A phylogene ic ee wi h all sub ype C conca ena ed agmen s om he NFLGs o CRF108_BC, and mos simila da abase sequences showed he closes ela ionship wi h he B azilian i us 02BR2022 om Sao Paulo (Figu e 4). Simila i y sea ches wi h he ou sub ype B agmen s and subsequen phylogene ic analyses wi h indi idual o conca ena ed agmen s ailed o iden i y any da abase i us ela ed o he sub ype B pa en al s ain o CRF108_BC. Vi uses 2021,13, 93 7 o 13 Vi uses 2021, 13, x FOR PEER REVIEW 8 o 14 Figu e 4. Phylogene ic ee o conca ena ed sub ype C agmen s o CRF108_BC. Only boo s ap alues ≥90% a e shown. CRF108_BC sequences a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he coun y o sample collec ion, i us name, and GenBank accession. 3.3. Phylogeog aphic and Phylodynamic Analyses To es ima e he empo al and geog aphic o igin o CRF108_BC, a Bayesian coalescen analysis was pe o med wi h he 15 PR-RT sequences o he C_2 clus e and 50 sub ype C da abase sequences om di e en coun ies o which yea and loca ion o sample collec- ion we e a ailable. P io o his analysis, he exis ence o a empo al signal was checked wi h TempEs 1.5.1, which e ealed a clock-like s uc u e in he da a se ( 2 = 0.38), indi- ca ing su icien empo al signal o pe o m he analyses. The MRCA o he clus e was es ima ed a ound 2000 (95% HPD, 1995–2004) in he ci y o Vi o ia wi h a loca ion pos e io p obabili y o 0.998. An ances y in B azil was also s ongly suppo ed (Figu e 5). Figu e 4. Phylogene ic ee o conca ena ed sub ype C agmen s o CRF108_BC. Only boo s ap alues ≥ 90% a e shown. CRF108_BC sequences a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO code o he coun y o sample collec ion, i us name, and GenBank accession. 3.3. Phylogeog aphic and Phylodynamic Analyses To es ima e he empo al and geog aphic o igin o CRF108_BC, a Bayesian coalescen analysis was pe o med wi h he 15 PR-RT sequences o he C_2 clus e and 50 sub ype C da abase sequences om di e en coun ies o which yea and loca ion o sample collec ion we e a ailable. P io o his analysis, he exis ence o a empo al signal was checked wi h TempEs 1.5.1, which e ealed a clock-like s uc u e in he da a se ( 2 = 0.38), indica ing su icien empo al signal o pe o m he analyses. The MRCA o he clus e was es ima ed a ound 2000 (95% HPD, 1995–2004) in he ci y o Vi o ia wi h a loca ion pos e io p obabili y o 0.998. An ances y in B azil was also s ongly suppo ed (Figu e 5). Vi uses 2021,13, 93 8 o 13 Vi uses 2021, 13, x FOR PEER REVIEW 9 o 14 Figu e 5. Maximum clade c edibili y ee o PR–RT sequences o CRF108_BC and 50 sub ype C sequences om da abases. Sequences belonging o he CRF108_BC clus e a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO coun y-code, i us name, and GenBank accession. Colo s o e minal and in e nal b anches ep esen sam- pling loca ions and mos p obable loca ions o he co esponding nodes, espec i ely, acco ding o he legend. Fo he nodes co esponding o he clus e and i s closes ances o , he mos p obable loca ions and he mean MRCA (wi h 95% HPD in e als) a e indica ed. Nodes suppo ed by PP (pos e io p obabili y) = 0.998–1 and PP = 0.95–0.9979 a e ma ked wi h illed and un illed ci cles, espec i ely. 29 b anches co esponding o sequences om Bo swana, Cyp us, E hiopia, Geo gia, Is ael, Kenya, Malawi, Senegal, Sou h A ica, Sweden, Tanzania, Uni ed Kingdom, Yemen, and Zambia we e collapsed o be e iewing. 4. Discussion The cha ac e iza ion o six HIV-1 NFLG sequences o BC ecombinan i uses ob- ained om epidemiologically-unlinked pa ien s showing a coinciden mosaic s uc u e di e en om p e iously iden i ied CRFs and clus e ing wi h a 100% boo s ap alue al- lowed o de ine a new CRF, designa ed CRF108_BC. Nine addi ional pa ial sequences a e g ouped in a monophyle ic clus e in PR-RT. All 15 pa ien s we e diagnosed wi h HIV-1 in ec ion in No he n Spain be ween 2006 and 2019 and we e in ec ed p edominan ly ia he e osexual con ac . CRF108_BC is he i s CRF ecombinan o sub ypes B and C pa en al s ains iden i- ied in Spain and he second in Eu ope (a e CRF60_BC in I aly [13]). Ten CRF_BCs ha e been iden i ied elsewhe e, nine in China [CRF07_BC [32], CRF08_BC [33], CRF57_BC [34], CRF61_BC [35], CRF62_BC [36], CRF64_BC [37], CRF85_BC [38], CRF86_BC [39] and CRF88_BC [40]] and one in B azil [CRF31_BC [41]]. Simila o CRF60_BC, he sub ype C pa en al s ain o CRF108_BC is phylogene ically ela ed o he sub ype C s ain ci cula - ing in B azil. Howe e , he ecombina ion e en gi ing ise o CRF108_BC could ha e Figu e 5. Maximum clade c edibili y ee o PR–RT sequences o CRF108_BC and 50 sub ype C sequences om da abases. Sequences belonging o he CRF108_BC clus e a e in blue and bold ype. Da abase sequences a e labeled wi h sub ype, wo-le e ISO coun y-code, i us name, and GenBank accession. Colo s o e minal and in e nal b anches ep esen sampling loca ions and mos p obable loca ions o he co esponding nodes, espec i ely, acco ding o he legend. Fo he nodes co esponding o he clus e and i s closes ances o , he mos p obable loca ions and he mean MRCA (wi h 95% HPD in e als) a e indica ed. Nodes suppo ed by PP (pos e io p obabili y) = 0.998–1 and PP = 0.95–0.9979 a e ma ked wi h illed and un illed ci cles, espec i ely. 29 b anches co esponding o sequences om Bo swana, Cyp us, E hiopia, Geo gia, Is ael, Kenya, Malawi, Senegal, Sou h A ica, Sweden, Tanzania, Uni ed Kingdom, Yemen, and Zambia we e collapsed o be e iewing. 4. Discussion The cha ac e iza ion o six HIV-1 NFLG sequences o BC ecombinan i uses ob ained om epidemiologically-unlinked pa ien s showing a coinciden mosaic s uc u e di e en om p e iously iden i ied CRFs and clus e ing wi h a 100% boo s ap alue allowed o de ine a new CRF, designa ed CRF108_BC. Nine addi ional pa ial sequences a e g ouped in a monophyle ic clus e in PR-RT. All 15 pa ien s we e diagnosed wi h HIV-1 in ec ion in No he n Spain be ween 2006 and 2019 and we e in ec ed p edominan ly ia he e osexual con ac . CRF108_BC is he i s CRF ecombinan o sub ypes B and C pa en al s ains iden i ied in Spain and he second in Eu ope (a e CRF60_BC in I aly [ 13 ]). Ten CRF_BCs ha e been iden i ied elsewhe e, nine in China [CRF07_BC [ 32 ], CRF08_BC [ 33 ], CRF57_BC [ 34 ], CRF61_BC [ 35 ], CRF62_BC [ 36 ], CRF64_BC [ 37 ], CRF85_BC [ 38 ], CRF86_BC [ 39 ] and CRF88_BC [ 40 ]] and one in B azil [CRF31_BC [ 41 ]]. Simila o CRF60_BC, he sub ype C pa en al s ain o CRF108_BC is phylogene ically ela ed o he sub ype C s ain ci cula ing in B azil. Howe e , he ecombina ion e en gi ing ise o CRF108_BC could ha e occu ed Vi uses 2021,13, 93 9 o 13 ei he in B azil, whe e bo h B and C sub ypes co-ci cula e a high p opo ions in some a eas [ 42 , 43 ] o in Spain, since we could no ack he ances y o he pa en al sub ype B s ain o any coun y. A Spanish o igin o CRF108_BC would be suppo ed by an in e ed Spanish MRCA in he ci y o Vi o ia, Basque Coun y, No he n Spain, a ound 2000. Howe e , we canno ule ou ha CRF108_BC could be ci cula ing a a low p e alence in some a ea(s) o B azil, which could explain i s lack o ep esen a ion in public sequence da abases. CRF108_BC, simila ly o all o he CRF_BCs iden i ied o da e, is p edominan ly o sub ype C. I is in e es ing o no e ha in all CRF_BCs, en is mos ly o sub ype C, which could lead o he specula ion ha a sub ype C en elope could con e e olu iona y ad an- ages ega ding i al i ness, escape o immune esponse o a mo e e icien eplica ion o ansmission o e a sub ype B en elope. Fu he in es iga ions a e equi ed o con i m any o hese possibili ies. To da e, he expansion o CRF108_BC has aken place om 2006 o 2019, wi h 53% o cases diagnosed in he las 4 yea s, and limi ed o a small geog aphical a ea in Spain, wi h 10 o 15 cases in he ci y o Vi o ia and 4 o he cases in wo neighbo ing p o inces o he Basque Coun y. Ou side o he Basque Coun y, we only ha e ound a single in ec ion wi h CRF108_BC in he ci y o Za agoza among HIV-1 sequences om mo e han 13,000 pa ien s om 10 egions o Spain analyzed by us and all sequences om public da abases. In Asia, he majo i y o CRF_BCs we e ansmi ed ini ially among pe sons who injec d ugs (PWID) [ 40 , 44 ] and ecen ly expanded among MSM [ 45 ]. In Eu ope, CRF60_BC is associa ed wi h p opaga ion among MSM [ 13 ]. Al hough, a p esen , MSM is he mos equen ansmission ou e o HIV-1 in Spain [ 46 ], wo CRFs p e iously iden i ied in Spain, CRF14_BG and CRF73_BG, p opaga ed mainly among PWID [ 7 , 15 ]. A hi d CRF iden i ied in Spain, CRF47_BF, was associa ed wi h he e osexual ansmission [ 10 ], simila ly o CRF108_BC. Howe e , we ha e obse ed u he p opaga ion o CRF47_BF among MSM [ 47 ]. This pa e n could po en ially be epea ed wi h CRF108_BC since, along wi h he e osexually ansmi ed cases (wi h 27% women), we ind an MSM diagnosed in 2017. This sugges s ha ansmission ne wo ks o HIV-1 in Spain among he e osexuals may be d i ing owa ds MSM. We ha e also obse ed he e e se si ua ion in he case o a CRF02_AG clus e sp eading om MSM o a he e osexual ne wo k [48]. NFLG sequencing o HIV-1 s ains in ol ed in expanding ansmission clus e s is highly ecommendable in o de o iden i y new CRFs, which may ha e acqui ed adap i e ad an ages h ough ecombina ion [ 49 , 50 ] e en when ecombina ion is no suspec ed in pa ial sequences, as i occu s wi h CRF108_BC. Gene ic di e si y and ecombina ion a e majo obs acles o he de elopmen o an e ec i e accine agains HIV-1 [ 51 – 53 ]. Cha ac e iza ion and molecula epidemiological su eillance o expanding new HIV-1 CRFs may play an impo an ole in public heal h ac ions, including he selec ion o op imal immunogens o e ec i e accines. 5. Conclusions A new HIV-1 CRF de i ed om sub ypes B and C, designa ed CRF108_BC, has been iden i ied a e he analysis o six NFLG om h ee ci ies in No he n Spain, which was o iginally iden i ied as a sub ype C clus e in PR–RT sequences comp ising 15 indi iduals. Phylogene ic and phylogeog aphic analyses poin o a B azilian ances y, al hough i is unclea whe he he ecombina ion e en ook place in B azil o in Spain. Among CRFs de i ed om B and C sub ypes, CRF108_BC is he 12 h iden i ied, he i s in Spain and he second in Eu ope. I s sp ead is cu en ly limi ed, wi h only 15 cases de ec ed, mos o hem ansmi ed ia he e osexual con ac . Conside ing ha mo e han hal o hem we e diagnosed in he las ou yea s, molecula epidemiological su eillance seems jus i ied o examine u he sp ead. The esul s o his s udy also ad oca e o NFLG sequence cha ac e iza ion o eme ging HIV-1 clus e s, which may ep esen new CRFs.