scieee Open visual document viewer

Word entropy-based approach to detect highly variable genetic markers for bacterial genotyping

Jakubíčková, Markéta; Bartoň, Vojtěch; Sedlář, Karel; Bezdíček, Matěj; Lengerová, Martina; Vítková, Helena

Abstract

Genotyping methods are used to distinguish bacterial strains from one species. Thus, distinguishing bacterial strains on a global scale, between countries or local districts in one country is possible. However, the highly selected bacterial populations (e.g. local populations in hospital) are typically closely related and low diversified. Therefore, currently used typing methods are not able to distinguish individual strains from each other. Here, we present a novel pipeline to detect highly variable genetic segments for genotyping a closely related bacterial population. The method is based on a degree of disorder in analyzed sequences that can be represented by sequence entropy. With the identified variable sequences, it is possible to find out transmission routes and sources of highly virulent and multiresistant strains. The proposed method can be used for any bacterial population, and due to its whole genome range, also noncoding regions are examined.

Full text

ORIGINAL RESEARCH published: 03 Feb ua y 2021 doi: 10.3389/ micb.2021.631605 F on ie s in Mic obiology | www. on ie sin.o g 1Feb ua y 2021 | Volume 12 | A icle 631605 Edi ed by: Dominique La enie , UMR6074 Ins i u de Reche che en In o ma ique e Sys èmes Aléa oi es (IRISA), F ance Re iewed by: Mahend a Ma iadassou, Ins i u Na ional de Reche che pou l’Ag icul u e, l’Alimen a ion e l’En i onnemen (INRAE), F ance William She win, Uni e si y o New Sou h Wales, Aus alia *Co espondence: Ma ke a Nyk yno a nyk yno a@ u b .cz Special y sec ion: This a icle was submi ed o E olu iona y and Genomic Mic obiology, a sec ion o he jou nal F on ie s in Mic obiology Recei ed: 20 No embe 2020 Accep ed: 13 Janua y 2021 Published: 03 Feb ua y 2021 Ci a ion: Nyk yno a M, Ba on V, Sedla K, Bezdicek M, Lenge o a M and Sku ko a H (2021) Wo d En opy-Based App oach o De ec Highly Va iable Gene ic Ma ke s o Bac e ial Geno yping. F on . Mic obiol. 12:631605. doi: 10.3389/ micb.2021.631605 Wo d En opy-Based App oach o De ec Highly Va iable Gene ic Ma ke s o Bac e ial Geno yping Ma ke a Nyk yno a1*, Voj ech Ba on1, Ka el Sedla 1, Ma ej Bezdicek2, Ma ina Lenge o a2and Helena Sku ko a1 1Depa men o Biomedical Enginee ing, Facul y o Elec ical Enginee ing and Communica ion, B no Uni e si y o Technology, B no, Czechia, 2Depa men o In e nal Medicine - Hema ology and Oncology, Uni e si y Hospi al B no, B no, Czechia Geno yping me hods a e used o dis inguish bac e ial s ains om one species. Thus, dis inguishing bac e ial s ains on a global scale, be ween coun ies o local dis ic s in one coun y is possible. Howe e , he highly selec ed bac e ial popula ions (e.g., local popula ions in hospi al) a e ypically closely ela ed and low di e si ied. The e o e, cu en ly used yping me hods a e no able o dis inguish indi idual s ains om each o he . He e, we p esen a no el pipeline o de ec highly a iable gene ic segmen s o geno yping a closely ela ed bac e ial popula ion. The me hod is based on a deg ee o diso de in analyzed sequences ha can be ep esen ed by sequence en opy. Wi h he iden i ied a iable sequences, i is possible o ind ou ansmission ou es and sou ces o highly i ulen and mul i esis an s ains. The p oposed me hod can be used o any bac e ial popula ion, and due o i s whole genome ange, also non-coding egions a e examined. Keywo ds: geno yping, en opy, gene ic ma ke s, closely ela ed bac e ia, MLST 1. INTRODUCTION Heal hca e-associa ed in ec ions (HAIs) a e a se ious wo ldwide h ea wi h signi ican impac on mo ali y and mo bidi y o hospi alized pa ien s (A e ian e al., 2016). They a e de ined as in ec ions de eloped in a hospi al o o he heal h ca e acili y ha i s appea 48 h o mo e a e hospi al admission, o wi hin 30 days a e ha ing ecei ed heal h ca e (Haque e al., 2018). The e o e, he abili y o iden i y he sou ce o he in ec ion and moni o he sp ead o disease is essen ial. To each ha goal, a p ocess called bac e ial yping is used (Ruppi sch, 2016) as i can dis inguish indi idual bac e ial s ains om one species. Howe e , apid labo a o y and compu a ional echniques’ esolu ion o yping is limi ed and he abili y o dis inguish s ains among a low di e si y hospi al bac e ial popula ion is p ac ically impossible. Me hods such as pulsed- ield gel elec opho esis (PFGE) (Schwa z and Can o , 1984), epe i i e sequence-based PCR ( ep-PCR) (Sku ko a e al., 2019), mul ilocus sequence yping (MLST) (Maiden e al., 1998), o mini-MLST (Bezdicek e al., 2019) a e used o cha ac e ize and in es iga e ou b eaks. As housekeeping genes used in MLST o mini-MLST may no ha e su icien a iabili y in some cases, he e is s ill a need o de elop new app oaches o iden i y a iable sequences. One o hem is mul ispace yping (MST) which uses sequences occu ing in junk DNA (in e genic egions and pseudogenes) as hey show highe a iabili y han coding genome pa s (Foucaul e al., 2005). Ano he app oach is single-locus sequence yping (SLST), whe e one a iable sequence iden i ied based on a genome mining app oach can ha e simila disc imina o y powe o MLST (Scholz e al., 2014). Howe e , i p ecise Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion analyses o he low di e si ied bac e ial popula ion a e equi ed, he only me hod ha can be used is whole-genome sequencing (WGS). Fo WGS yping, wo main app oaches exis : single nucleo ide a ian analysis (SNV) and co e genome MLST (cgMLST) o whole genome MLST (wgMLST) (Hen i e al., 2017; Schü ch e al., 2018), whe e genes common o bac e ial isola es a e compa ed. Ne e heless, he yping me hods which equi e WGS a e no sui able o ou ine clinical p ac ice. Addi ionally, he compa ison o di e en s udies is di icul as di e en da a quali y assessmen , genome assembly, and esul s analysis a e done (Ruppi sch, 2016). In his a icle, we p esen a new app oach o localize highly a iable sequences in bac e ial genomes o Esche ichia coli and En e ococcus aecium. De ec ing new yping agmen s is based on a calcula ion o wo d en opy in genomes assembled om NGS da a. Po en ial a iable sequences a e e alua ed by phylogene ic analysis, and he ob ained esul s a e compa ed o cgMLST. The co e genome was chosen as he iden i ied ma ke s mus be p esen ed in all bac e ial s ains o gi en species. Fo ha eason, i is no app op ia e o s udy he pan-genome as i con ains he co e genome and he dispensable genome (Medini e al., 2005). The iden i ied ma ke s will be used o closely ela ed bac e ial yping ia s anda d labo a o y me hods wi hou he need o u he WGS. Ou app oach can be used o any bac e ia; he only equi emen is aligned isola e genomes. 2. MATERIALS AND METHODS 2.1. Da ase Fo his a icle, samples we e collec ed in he Uni e si y Hospi al o B no. In o al, 23 E. aecium isola es we e ob ained om se en hospi al depa men s be ween 6/2017 and 2/2018 and 21 E. coli isola es we e collec ed om a single hospi al depa men be ween 5/2019 and 7/2019. The e o e, he da ase s ha e di e en a iabili y a es as hey we e collec ed du ing di e en pe iods and om a di e en numbe o depa men s. Sequencing lib a ies we e p epa ed wi h KAPA Hype Plus Ki s (Roche, Swi ze land), and a quali y check was pe o med using a 2100 Bioanalyze (Agilen Technologies, USA). Lib a ies we e quan i ied wi h KAPA Lib a y Quan i ica ion ki (Roche, Swi ze land). Sequencing was pe o med on an Illumina MiSeq pla o m using MiSeq Reagen Ki 2 (500-cycles) (Illumina, USA) and pai ed-end eads (250 bp) we e acqui ed. 2.2. Re e ence-Based Genome Mapping A e he sequenced da a’s quali y con ol, he eads we e mapped ia BBMap ( 38.71, Bushnell, 2014) o he human genome (GRCh38.p13) o emo e con amina ion which could eme ge du ing sample p epa a ion. Reads which did no map o he human genome we e u he analyzed and quali y assessmen was done. Adap e s and low-quali y pa s o eads we e immed ia T immoma ic ( 0.36, Bolge e al., 2014). The eads leng h was abou 250 bp; hus he e e ence-based genome mapping app oach was chosen, and o his pu pose, he Bu ows-Wheele Alignmen MEM ( 0.7.17- 1188, Li, 2013) was used. The e e ence sequences we e ob ained om he GenBank da abase (Cla k e al., 2016). As e e ence sequence o assembly o E. aecium genomes was chosen CP003351.1 (Lam e al., 2012) and E. coli genomes we e assembled agains he sequence BA000007.3 (Hayashi, 2001). Then Sam ools ( 1.7, Li e al., 2009) was used o emo e low-quali y and duplica ed eads, and consensus sequences we e gene a ed. 2.3. The En opy-Based De ec ion o Va iable Sequences To loca e highly a iable egions in sequences, he en opy alue can be used. The DNA sequences en opy es ima ion (Schmi and He zel, 1997) is based on he Shannon-en opy (Shannon, 1948). In his pape , he en opy calcula ion modi ica ion as in e se alue o in o ma ion con en is u ilized. This modi ica ion commonly used o sequence logo de e mina ion (Schneide and S ephens, 1990) measu es he conse a ion o posi ion in bi s which be e espec s di e en cha se s wi h a iable con en o ambi alen bases. En opy Hi(Schneide and S ephens, 1990) can be calcula ed as Hi= − X a=A,C,G,T a,i·log2 a,i, (1) whe e iis he posi ion in he sequence, is he equency o one o ou nucleo ides’ occu ence (a= A, C, G, T) in posi ion iand can be calcula ed as a,i=na n, (2) whe e nais numbe o occu ences o one nucleo ide in one posi ion and nis numbe o all cha ac e s a he gi en posi ion. In he ideal case, we wan o ind such a egion capable o dis inguishing each isola e. Fo ha eason, he cu en me hod o en opy calcula ion in one posi ion (Nyk yno a e al., 2019) in he sequence’s alignmen canno be used, as i can only dis inguish ou a ian s because ou nucleo ides (A, C, G, T) a e p esen ed. In his a icle, a new unique app oach using wo d en opy is p esen ed, whe e he wo d is a nucleo ide sequence o de ined leng h. As e e ence-based genome assembly is no pe ec du ing consensus calling ambiguous nucleo ides can appea in he consensus sequences, and hey a i icially inc ease he en opy alue as shown in Figu e 1. To p e en de ec ing alsely a iable sequences, he whole columns o aligned sequences whe e ambiguous nucleo ides occu we e emo ed and o p ese e he o iginal leng h o he c ea ed consensus sequences eplaced by gaps. Also, he de ec ed ma ke s should occu in all analyzed isola es. Thus, when gaps appea in some sequences, he gaps ep esen ed by dashes a e added o he whole columns as is also shown in Figu e 1. In he window w, whose size co esponds o wo d leng h, he numbe o unique wo ds is de e mined, and hen he en opy o wo ds Hwis calcula ed as Hw= − X(kw n·(log2 kw n)), (3) F on ie s in Mic obiology | www. on ie sin.o g 2Feb ua y 2021 | Volume 12 | A icle 631605 Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion FIGURE 1 | Example o sequence p ep ocessing. (A) Example o aligned sequences be o e p ep ocessing and hei en opy o wo d alues o wo ds o leng h i e. (B) Aligned sequences a e p ep ocessing and co esponding en opy o wo d alues. FIGURE 2 | Scheme o en opy o wo ds’ calcula ion o a se o sequences a e p ep ocessing and o wo d leng h o 5 bp and one nucleo ide window shi . whe e kwis he numbe o occu ences o unique wo ds in each alignmen in he cu en window, and nis he numbe o isola es. The sum is applied o e he alignmen in he cu en window. A e es ima ing he en opy alue, he window mo es wi h a shi o one nucleo ide. The p inciple o en opy o wo ds’ calcula ion is shown in Figu e 2. Fo de ined sliding window sizes w(also called wo d leng h), he en opy signal wi h leng h N−w+1 is ob ained, whe e Nis he analyzed genome’s leng h. In he signal, he maximum alue is iden i ied. In heo y, he maximum alue can be de ined as a bina y loga i hm o he numbe o sequences. Based on he maximum alue F on ie s in Mic obiology | www. on ie sin.o g 3Feb ua y 2021 | Volume 12 | A icle 631605 Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion ha occu s in he signal, he h eshold is se as a maximum alue educed by 0.4 bi , and signal peaks abo e he h eshold a e labeled as po en ial gene ic ma ke s. The s a ed alue was chosen empi ically based on he numbe o po en ial a iable egions o ensu e ha some egions will ha e su icien disc imina o y powe ; howe e , i can be adjus ed o analyzed da a by he use . The aligned sequences co esponding o peaks a e ex ac ed and analyzed. I wo peaks a e dis an less han he window leng h, he sequences co esponding o he peaks a e me ged and analyzed as one. De eloped sou ce codes and unc ions in Ma lab R2020a can be ound a h ps://gi hub.com/ma ke anyk yno a/ma ke s_de ec ion. 2.4. Phylogene ic Analysis In he i s pa , each se o aligned sequences co esponding o peaks was analyzed. The e olu iona y dis ances we e calcula ed based on he Kimu a (Kimu a, 1980) model, which is used as s anda d. F om ob ained dis ances, he phylogene ic ees we e cons uc ed based on UPGMA. Clus e analysis was conduc ed, and he clus e s om he ees we e es ablished and compa ed o esul s ob ained om cgMLST. The h eshold o clus e s de e mina ion was se o in e sec he ee so ha he maximum numbe o clus e s wasob ained. The esul s showed ha no se had su icien disc imina o y powe o dis inguish be ween s ains. Tha is why he combina ion o wo aligned sequences se s was used. The iden i ied a iable se s o sequences we e me ged, which c ea ed one long sequence o each analyzed bac e ial isola e. Then o me ged sequences, again he e olu iona y dis ances we e calcula ed and based on hem, he phylogene ic ees we e cons uc ed and examined. The leng h o a a iable se o sequences had a p opo ional impac on he c ea ed phylogene ic ee. I he sequences in one a iable se had a longe leng h, hey had a highe impac on he ee and ice e sa, he se o sho e sequences did no a ec he inal ee so much. Ne e heless, sequences a iabili y has he main impac on he phylogene ic ee. 3. RESULTS AND DISCUSSION 3.1. Sequence Type De e mina ion and cgMLST Analysis The inco po a ed E. coli and E. aecium cgMLST, cgSNV, and wgSNV schemes in Ridom SeqSphe e+ so wa e (Ridom, DE) we e used o analyze he gene ic ela edness o sequenced bac e ial genomes. Fo E. coli, his included 3,152 a ge genes used o gene a e a cgMLST dend og am. Fo he cgSNV minimum spanning ee analysis, 152,212 aligned nucleo ide si es we e analyzed. In o al, 176,629 aligned nucleo ide si es we e used o wgSNV minimum spanning ee analysis. Fo E. aecium, he scheme included 1,423 a ge genes o cgMLST, 7,707 aligned nucleo ide si es o cgSNV and 10,858 aligned nucleo ide si es o wgSNV analysis. The WGS da a we e used o de e mine sequence ypes (STs) o bo h, E. coli and E. aecium s ains. Fo E. coli, se en housekeeping genes (adk, umC, gy B, lcd, mdh, pu A, ecA) and o E. aecium again se en housekeeping genes (a pA, ddl, gdh, pu K, gyd, ps S, adk) we e analyzed. Fo E. aecium h ee STs we e iden i ied (4 ×ST 117, 1 ×ST 18, 18 ×ST 80), and o E. coli ele en STs we e ecognized and one ST emains unknown (1 ×ST 69, 4 ×ST 131, 1 ×ST 95, 2 ×ST 404, 2 ×ST 38, 2 ×ST 1049, 4 ×ST 58, 1 ×ST 297, 1 ×ST 517, 2 ×ST 101, 1 ×ST UNW). The comple e esul s a e a ached as Supplemen a y File 1. 3.2. Po en ial Gene ic Ma ke s Analysis Sequences wi h a high alue o wo d en opy we e iden i ied o wo d leng hs om 50 o 400 bp wi h a s ep o 50 bp. T adi ional labo a o y me hods can use iden i ied ma ke s o leng h om his ange. In he i s pa , one a iable agmen wi h high wo d en opy alues was used o dis inguish he s ains o E. coli and E. aecium. Examining 1,358 agmen s wi h high a iabili y o all wo d sizes showed ha he maximum numbe o clus e s ha can be dis inguished o E. coli was 13. Fo E. aecium, 104 po en ial gene ic ma ke s o all wo d sizes we e analyzed, and he maximum numbe o dis inguishable clus e s was 5. The comple e esul s a e in Supplemen a y File 2. The numbe o dis inguishable clus e s using one a iable agmen is no high enough. Fo ha eason, wo a iable agmen s wi h wo d en opy abo e h eshold we e analyzed. Fo all combina ions o wo possible gene ic ma ke s, he phylogene ic ees we e cons uc ed, and he numbe o clus e s was calcula ed. In he nex s ep, he po en ial gene ic ma ke s wi h he highes numbe o clus e s we e analyzed. Fo all wo d sizes, 136,661 combina ions o possible gene ic ma ke s o E. coli and 685 o E. aecium we e examined in o al. In case o need, he combina ion o e en mo e ma ke s can be used; howe e , he ime o analysis will g ow as he ime complexi y is O(nm), whe e mis he numbe o a iable agmen s. Fo E. coli, 29 ees we e examined which di ided he isola es o 15 clus e s. Ten ees classi y he genomes acco ding o he 11 p ede ined clus e s om cgMLST. The esul s om cgMLST wi h 11 ma ked sequence ypes and 11 de ined clus e s and a cladog am c ea ed based on he combina ion o wo gene ic ma ke s wi h 15 labeled clus e s a e shown in Figu e 3. In each case o one gene ic ma ke , he egion which con ained he genes o hypo he ical p o ein (NP_308436.2) and S- o mylglu a hione hyd olase (NP_308437.1) was selec ed. In i e ou o en cases, he in e genic egion in on o hypo he ical p o ein was also inco po a ed. As he second gene ic ma ke , wo egions we e mos o en ( h ee imes) chosen. The i s one con ained ou e memb ane p o ein (NP_313225.1) and in one case he in e genic egion in on o he gene oo. The second one s a ed in imb ial-like adhesin p o ein (NP_310944.1) and ended in hypo he ical p o ein (NP_310945.1). Fou o he gene ic ma ke s we e selec ed only one ime. Two o hem consis ed o non-coding egion and gene [PhoH amily P-loop ATPase (NP_309293.1) in he i s case, excinuclease U ABC subuni U C (NP_310678.2) in he second case], o ano he ma ke pa o he gene alyl- RNA syn he ase (NP_313262.1) was picked, and in he las case, as a gene ic ma ke he non- coding egion was selec ed. De ailed in o ma ion is a ailable in F on ie s in Mic obiology | www. on ie sin.o g 4Feb ua y 2021 | Volume 12 | A icle 631605 Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion FIGURE 3 | Classi ica ion o 21 E. coli genomes. (A) Co e genome MLST analysis o 21 E. coli isola es wi h p ede ined ma ked clus e s (ou e ci cles) and sequence ypes (inne ci cles). (B) Cladog am o 21 E. coli isola es based on wo iden i ied gene ic ma ke s wi h highligh clus e s ob ained om co e genome MLST analysis (ou e ci cle) and subg oups ob ained based on analysis o iden i ied a iable sequences (inne ci cle) c ea ed by E ol iew (Sub amanian e al., 2019). FIGURE 4 | Classi ica ion o 23 E. aecium genomes. (A) Co e genome MLST analysis o 23 E. aecium isola es wi h p ede ined ma ked clus e s (ou e ci cles) and sequence ypes (inne ci cles). (B) Cladog am o 23 E. aecium isola es based on wo iden i ied gene ic ma ke s wi h highligh clus e s ob ained om co e genome MLST analysis (ou e ci cle) and subg oups ob ained based on analysis o iden i ied a iable sequences (inne ci cle) c ea ed by E ol iew (Sub amanian e al., 2019). Supplemen a y File 2 and all phylogene ic ees ha co ec ly classi ied genomes a e depic ed in Supplemen a y Figu e 1. Fo E. aecium, 22 ees which di ided he isola es in o nine clus e s we e examined. I was ound ha six ees classi y he genomes acco ding o i e p ede ined clus e s om cgMLST. In Figu e 4, he esul s om cgMLST analysis wi h h ee labeled sequence ypes, i e clus e s and a cladog am cons uc ed based on he iden i ied combina ion o wo gene ic ma ke s wi h nine labeled clus e s shown. As in he p e ious case, in all possible combina ions, one gene ic ma ke always emained he same, and i was RNA me hyl ans e ase (WP_002294889.1). In one o six cases, he in e genic egion a ound he gene was also inco po a ed in o he a iable gene ic pa . Mg- ansloca ing P- ype ATPase (WP_002304253.1) was in mos cases ( h ee imes) selec ed as he second ma ke . The egion ha con ained he gene o DUF1430 domain-con aining p o ein (WP_002288541.1) was chosen wice. Fo he las possible gene ic ma ke , he gene o nucleoside-diphospha e kinase (WP_002293327.1) was picked. The comple e esul s a e a ailable in Supplemen a y File 2 and all phylogene ic ees co ec ly classi ying genomes a e shown in Supplemen a y Figu e 2. 3.3. Wo d Sizes Fo de ined wo d sizes, a maximum numbe o co ec ly classi ied clus e s o phylogene ic ees based on wo ma ke combina ions F on ie s in Mic obiology | www. on ie sin.o g 5Feb ua y 2021 | Volume 12 | A icle 631605 Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion FIGURE 5 | Analysis o wo d size impac on numbe o co ec ly classi ied clus e s. (A) Maximum numbe o co ec ly classi ied clus e s o combina ions o wo a iable agmen s o di e en wo d sizes and numbe o wo a iable agmen combina ions ha co ec ed classi ied analyzed E. aecium isola es. (B) Maximum numbe o co ec ly classi ied clus e s o combina ion o wo a iable agmen s o di e en wo d sizes and numbe o wo a iable agmen combina ions ha co ec ed classi ied analyzed E. coli isola es. was es ablished, as is shown in Figu e 5. Fo analyzed ees he numbe o clus e s was de e mined and he co ec genomes classi ica ion was con olled. The esul s showed ha o E. aecium he wo d leng hs do no signi ican ly a ec he maximum numbe o co ec ly es ablished clus e s, which is nine o all wo d sizes excep he leng h 50 and 300 bp. On he con a y, he esolu ion powe o E. coli g ows wi h inc easing wo d leng hs and he bes esul s (dis inguish 15 clus e s) can be ob ained o a window o size 300 bp and mo e. 3.4. Resul s Ve i ica ion The analyzed da ase s we e ex ended by comple e genomes sequences o E. coli and E. aecium ob ained om he GenBank da abase. Fo each bac e ium, six sequences we e downloaded, and he STs we e de e mined. Fo E. coli isola es (CP050212.1, CP050219.1, CP050211.1, CP050201.1, CP050214.1, CP050205.1) i e STs we e iden i ied (1 ×ST 2705, 1×ST 2280, 1 ×ST 70, 1 ×ST 602, 2 ×ST 131), and o E. aecium isola es (CP040236.1, CP036151.1, CP027501.1, CP035666.1, CP035660.1, CP035220.1) one ST was ecognized, and one ST emains unknown (5 ×ST 80, 1 ×ST UNW). The esul s a e in Supplemen a y File 1. Using he BLAST+ ( 2.6.0+, Camacho e al., 2009) he iden i ied a iable ma ke s we e loca ed in hese genomes. The ma ke s sequences o bo h se s (o iginal and da abase) we e analyzed oge he . The sequences we e aligned, e olu iona y dis ances we e calcula ed, and phylogene ic ees we e cons uc ed. To ally, 27 isola es o E. coli belonging o nine STs we e classi ied o 17 clus e s and 29 isola es o E. aecium belonging o ou STs we e i e imes classi ied o 11 clus e s and once o 10 clus e s. All c ea ed phylogene ic ees a e depic ed in Supplemen a y Figu e 3. The iden i ied gene ic ma ke s can be used o yping no el isola es. Howe e , i a la ge numbe o new isola es should be analyzed, he algo i hm should be un again o ind ma ke s wi h he highes disc imina o y powe o he analyzed bac e ial popula ion. 3.5. Analyzed Bac e ia The p oposed app oach was es ed on wo bac e ia—E. aecium and E. coli which belong among signi ican HAIs pa hogens. E. aecium is a G am-posi i e bac e ium and can be a sou ce o nosocomial in ec ions, especially in immunocomp omised pa ien s, and causes endoca di is, bac e emia o u ina y ac in ec ion. In addi ion, some E. aecium isola es can de elop d ug esis ance o se e al an ibio ics g oups such as glycopep ides, be a-lac ams, luo oquinolones, o aminoglycosides (Yoong e al., 2004; Cas illo-Rojas e al., 2013). Ano he pa hogen om G am-nega i e bac e ia is E. coli, which can be ound in human in es inal lo a as a commensal. Howe e , as an oppo unis ic pa hogen E. coli ep esen s a huge public heal h p oblem. I is one o he mos common causes o HAIs ypically connec ed o meningi is, pneumonia, u ina y ac in ec ions, and so issue in ec ions (Jau eguy e al., 2008). E. coli s ains ha e he abili y o accumula e esis ance genes; hus, he esis ance o b oad-spec um cephalospo ins, ca bapenems, aminoglycoside, luo oquinolones, and polymyxins is o en obse ed (Poi el e al., 2018). F on ie s in Mic obiology | www. on ie sin.o g 6Feb ua y 2021 | Volume 12 | A icle 631605 Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion 4. CONCLUSION As yping is c ucial o iden i ying he sou ce o in ec ion and ou b eak moni o ing, i is necessa y o ha e a highly sensi i e yping me hod o dis inguish closely ela ed bac e ial s ains. Howe e , cu en ly used me hods ha e insu icien disc imina o y powe excep whole genome sequencing. Ne e heless, WGS is ime-consuming, and da a p ocessing is no possible in ou ine clinical p ac ice. Fo his eason, a new app oach o de ec a iable gene ic ma ke s om WGS da a is p esen ed. The WGS is only used once o ob ain he i s se o da a o gene ic ma ke s de ec ion. To loca e highly a iable sequences, wo d en opy is employed. The loca ed ma ke s can be used o yping a closely ela ed bac e ial popula ion ia s anda d labo a o y me hods. Thus, ano he genome sequencing will be necessa y only i a la ge numbe o new isola es should be analyzed. Using he en opy-based app oach, new gene ic ma ke s wi h he same o e en highe disc imina o y powe han cgMLST can be iden i ied. As he leng h o wo ds om 50 o 400 bp wi h a s ep o 50 bp is applied, po en ial ma ke s can be used in s anda d labo a o y echniques because hey ha e sui able leng h. The only equi emen o he p oposed app oach is ha genomes mus be aligned, which can cause a p oblem mainly when la ge genomes should be examined. The p esen ed me hod was es ed on bo h G am-posi i e and G am-nega i e bac e ia. As in bo h es ed da ase s, one ma ke wi h su icien disc imina o y powe o yping o low di e si y bac e ial popula ion was no ound, he combina ion o wo ma ke s was employed. Based on wo iden i ied ma ke s, we we e able o dis inguish he G am-nega i e and G am-posi i e bac e ial popula ion a he same and e en a highe disc imina ion le el han ia classic labo a o y me hods using only wo a iable sequences ins ead o se en housekeeping genes. The iden i ied ma ke s will be used in ou ine labo a o y me hods o bac e ial yping. The low numbe o a iable agmen s ha should be examined will sho en he ime o analysis. Commonly used yping me hods mainly use coding egions and analyze genes conse ed and p esen ed in all isola es o he gi en species. Hence he a iabili y a e is lowe and yping he closely ela ed popula ion is no possible. The p oposed me hod uses he whole genome ange; hus, also in e genic genome egions a e analyzed as hey ha e highe a iabili y han coding a eas and can con ain gene ic ma ke s o yping ela ed isola es. Based on iden i ied ma ke s, he closely ela ed bac e ial popula ion can be di e si ied and o ind ou he ansmission ou es and sou ce o ou b eaks will be possible. DATA AVAILABILITY STATEMENT The da ase s p esen ed in his s udy can be ound in online eposi o ies. The names o he eposi o y/ eposi o ies and accession numbe (s) can be ound a : h ps://www.ncbi.nlm.nih. go /, PRJNA675431. AUTHOR CONTRIBUTIONS MN, VB, KS, and HS con ibu ed o he concep ion and design o he s udy. MN implemen ed he algo i hm and e alua ed he esul s. ML and MB ensu ed he biological aspec s o he p ojec . All au ho s ead and app o ed he inal manusc ip . FUNDING This wo k was suppo ed by a g an p ojec om he Czech Science Founda ion (GA17-01821S). SUPPLEMENTARY MATERIAL The Supplemen a y Ma e ial o his a icle can be ound online a : h ps://www. on ie sin.o g/a icles/10.3389/ micb. 2021.631605/ ull#supplemen a y-ma e ial Supplemen a y File 1 | MLST analysis esul s o E.coli and E. aecium. Supplemen a y File 2 | Po en ial and inal gene ic ma ke s o E. coli and E. aecium. Supplemen a y Figu e 1 | Phylogene ic ees ha co ec ly classi ied E. coli genomes wi h ma ked clus e s. Supplemen a y Figu e 2 | Phylogene ic ees ha co ec ly classi ied E. aecium genomes wi h ma ked clus e s. Supplemen a y Figu e 3 | Phylogene ic ees o o iginal and da abase genomes o E. aecium and E. coli genomes wi h ma ked clus e s. REFERENCES A e ian, H., Vogel, M., Kwe ka , A., and Ha mann, M. (2016). Economic e alua ion o in e en ions o p e en ion o hospi al acqui ed in ec ions: a sys ema ic e iew. PLoS ONE 11:e0146381. doi: 10.1371/jou nal.pone.0146381 Bezdicek, M., Nyk yno a, M., Ple o a, K., B helo a, E., Kocmano a, I., Sedla , K., e al. (2019). Applica ion o mini-MLST and whole genome sequencing in low di e si y hospi al ex ended-spec um be a-lac amase p oducing Klebsiella pneumoniae popula ion. PLoS ONE 14:e0221187. doi: 10.1371/jou nal.pone.0221187 Bolge , A. M., Lohse, M., and Usadel, B. (2014). T immoma ic: a lexible imme o Illumina sequence da a. Bioin o ma ics 30, 2114–2120. doi: 10.1093/bioin o ma ics/b u170 Bushnell, B. (2014). BBMap: A Fas , Accu a e, Splice-Awa e Aligne . No. LBNL- 7065E. Be keley, CA: E nes O lando Law ence Be keley Na ional Labo a o y. Camacho, C., Coulou is, G., A agyan, V., Ma, N., Papadopoulos, J., Beale , K., e al. (2009). BLAST+: a chi ec u e and applica ions. BMC Bioin o ma ics 10:421. doi: 10.1186/1471-2105-10-421 Cas illo-Rojas, G., Maza i-Hi ía , M., Ponce de León, S., Amie a-Fe nández, R. I., Agis-Juá ez, R. A., Huebne , J., e al. (2013). Compa ison o En e ococcus aecium and En e ococcus aecalis s ains isola ed om wa e and clinical samples: an imic obial suscep ibili y and gene ic ela ionships. PLoS ONE 8:e59491. doi: 10.1371/jou nal.pone. 0059491 Cla k, K., Ka sch-Miz achi, I., Lipman, D. J., Os ell, J., and Saye s, E. W. (2016). GenBank. Nucleic Acids Res. 44, D67–D72. doi: 10.1093/na /gk 1276 Foucaul , C., La Scola, B., Lind oos, H., Ande sson, S. G. E., and Raoul , D. (2005). Mul ispace yping echnique o sequence-based yping o Ba onella quin ana. J. Clin. Mic obiol. 43, 41–48. doi: 10.1128/JCM.43.1.41-4 8.2005 F on ie s in Mic obiology | www. on ie sin.o g 7Feb ua y 2021 | Volume 12 | A icle 631605 Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion Haque, M., Sa elli, M., McKimm, J., and Abu Baka , M. B. (2018). Heal h ca e-associa ed in ec ions-an o e iew. In ec . D ug Resis . 11, 2321–2333. doi: 10.2147/IDR.S177247 Hayashi, T. (2001). Comple e genome sequence o en e ohemo hagic Eschelichia coli O157:H7 and genomic compa ison wi h a labo a o y s ain K-12. DNA Res. 8, 11–22. doi: 10.1093/dna es/8.1.11 Hen i, C., Leeki cha oenphon, P., Ca le on, H. A., Radomski, N., Kaas, R. S., Ma ie , J.-F., e al. (2017). An assessmen o di e en genomic app oaches o in e ing phylogeny o Lis e ia monocy ogenes. F on . Mic obiol. 8:2351. doi: 10.3389/ micb.2017.02351 Jau eguy, F., Land aud, L., Passe , V., Diancou , L., F apy, E., Guigon, G., e al. (2008). Phylogene ic and genomic di e si y o human bac e emic Esche ichia coli s ains. BMC Genomics 9:560. doi: 10.1186/1471-2164-9-560 Kimu a, M. (1980). A simple me hod o es ima ing e olu iona y a es o base subs i u ions h ough compa a i e s udies o nucleo ide sequences. J. Mol. E ol. 16, 111–120. doi: 10.1007/BF01731581 Lam, M. M. C., Seemann, T., Bulach, D. M., Gladman, S. L., Chen, H., Ha ing, V., e al. (2012). Compa a i e analysis o he i s comple e En e ococcus aecium genome. J. Bac e iol. 194, 2334–2341. doi: 10.1128/JB.00259-12 Li, H. (2013). Aligning sequence eads, clone sequences and assembly con igs wi h BWA-MEM. a Xi [P ep in ]. a Xi :1303.3997 Li, H., Handsake , B., Wysoke , A., Fennell, T., Ruan, J., Home , N., e al. (2009). The Sequence Alignmen /Map o ma and SAM ools. Bioin o ma ics 25, 2078–2079. doi: 10.1093/bioin o ma ics/b p352 Maiden, M. C. J., Byg a es, J. A., Feil, E., Mo elli, G., Russell, J. E., U win, R., e al. (1998). Mul ilocus sequence yping: a po able app oach o he iden i ica ion o clones wi hin popula ions o pa hogenic mic oo ganisms. P oc. Na l. Acad. Sci. U.S.A. 95, 3140–3145. doi: 10.1073/pnas.95.6.3140 Medini, D., Dona i, C., Te elin, H., Masignani, V., and Rappuoli, R. (2005). The mic obial pan-genome. Cu . Opin. Gene . De . 15, 589–594. doi: 10.1016/j.gde.2005.09.006 Nyk yno a, M., Made anko a, D., Ba on, V., Bezdicek, M., Lenge o a, M., and Sku ko a, H. (2019). “En opy-based de ec ion o gene ic ma ke s o bac e ia geno yping,” in Lec u e No es in Compu e Science (including subse ies Lec u e No es in A i icial In elligence and Lec u e No es in Bioin o ma ics) (Cham: Sp inge ), 11466 LNBI, 177–188. doi: 10.1007/978-3-030-17935-9_17 Poi el, L., Madec, J.-Y., Lupo, A., Schink, A.-K., Kie e , N., No dmann, P., e al. (2018). “An imic obial esis ance in Esche ichia coli,” in An imic obial Resis ance in Bac e ia om Li es ock and Companion Animals, eds S. Schwa z, L. M. Ca aco, and J. Shen (Washing on, DC: Ame ican Socie y o Mic obiology), 289–316. doi: 10.1128/mic obiolspec.ARBA-0026-2017 Ruppi sch, W. (2016). Molecula yping o bac e ia o epidemiological su eillance and ou b eak in es iga ion/Molekula e Typisie ung on Bak e ien ü die epidemiologische Übe wachung und Ausb uchsabklä ung. J. Land Manage. Food En i on. 67, 199–224. doi: 10.1515/boku-20 16-0017 Schmi , A. O., and He zel, H. (1997). Es ima ing he en opy o DNA sequences. J. Theo . Biol. 188, 369–377. doi: 10.1006/j bi.1997.0493 Schneide , T. D., and S ephens, R. M. (1990). Sequence logos: a new way o display consensus sequences. Nucl. Acids Res. 18, 6097–6100. doi: 10.1093/na /18.20.6097 Scholz, C. F. P., Jensen, A., Lomhol , H. B., B üggemann, H., and Kilian, M. (2014). A no el high- esolu ion single locus sequence yping scheme o mixed popula ions o P opionibac e ium acnes in i o. PLoS ONE 9:e104199. doi: 10.1371/jou nal.pone.0104199 Schü ch, A., A edondo-Alonso, S., Willems, R., and Goe ing, R. (2018). Whole genome sequencing op ions o bac e ial s ain yping and epidemiologic analysis based on single nucleo ide polymo phism e sus gene-by-gene-based app oaches. Clin. Mic obiol. In ec . 24, 350–354. doi: 10.1016/j.cmi.2017.12.016 Schwa z, D. C., and Can o , C. R. (1984). Sepa a ion o yeas ch omosome- sized DNAs by pulsed ield g adien gel elec opho esis. Cell 37, 67–75. doi: 10.1016/0092-8674(84)90301-5 Shannon, C. E. (1948). A ma hema ical heo y o communica ion. Bell Sys . Techn. J. 27, 379–423. doi: 10.1002/j.1538-7305.1948. b01338.x Sku ko a, H., Vi ek, M., Bezdicek, M., B helo a, E., and Lenge o a, M. (2019). Ad anced DNA inge p in geno yping based on a model de eloped om eal chip elec opho esis da a. J. Ad . Res. 18, 9–18. doi: 10.1016/j.ja e.2019.01.005 Sub amanian, B., Gao, S., Le che , M. J., Hu, S., and Chen, W.-H. (2019). E ol iew 3: a webse e o isualiza ion, anno a ion, and managemen o phylogene ic ees. Nucl. Acids Res. 47, W270-W275. doi: 10.1093/na /gkz357 Yoong, P., Schuch, R., Nelson, D., and Fische i, V. A. (2004). Iden i ica ion o a b oadly ac i e phage ly ic enzyme wi h le hal ac i i y agains an ibio ic- esis an En e ococcus aecalis and En e ococcus aecium. J. Bac e iol. 186, 4808–4812. doi: 10.1128/JB.186.14.4808-4812.2004 Con lic o In e es : The au ho s decla e ha he esea ch was conduc ed in he absence o any comme cial o inancial ela ionships ha could be cons ued as a po en ial con lic o in e es . Copy igh © 2021 Nyk yno a, Ba on, Sedla , Bezdicek, Lenge o a and Sku ko a. This is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons A ibu ion License (CC BY). The use, dis ibu ion o ep oduc ion in o he o ums is pe mi ed, p o ided he o iginal au ho (s) and he copy igh owne (s) a e c edi ed and ha he o iginal publica ion in his jou nal is ci ed, in acco dance wi h accep ed academic p ac ice. No use, dis ibu ion o ep oduc ion is pe mi ed which does no comply wi h hese e ms. F on ie s in Mic obiology | www. on ie sin.o g 8Feb ua y 2021 | Volume 12 | A icle 631605