ORIGINAL RESEARCH
published: 03 Feb ua y 2021
doi: 10.3389/ micb.2021.631605
F on ie s in Mic obiology | www. on ie sin.o g 1Feb ua y 2021 | Volume 12 | A icle 631605
Edi ed by:
Dominique La enie ,
UMR6074 Ins i u de Reche che en
In o ma ique e Sys èmes Aléa oi es
(IRISA), F ance
Re iewed by:
Mahend a Ma iadassou,
Ins i u Na ional de Reche che pou
l’Ag icul u e, l’Alimen a ion e
l’En i onnemen (INRAE), F ance
William She win,
Uni e si y o New Sou h Wales,
Aus alia
*Co espondence:
Ma ke a Nyk yno a
nyk yno a@ u b .cz
Special y sec ion:
This a icle was submi ed o
E olu iona y and Genomic
Mic obiology,
a sec ion o he jou nal
F on ie s in Mic obiology
Recei ed: 20 No embe 2020
Accep ed: 13 Janua y 2021
Published: 03 Feb ua y 2021
Ci a ion:
Nyk yno a M, Ba on V, Sedla K,
Bezdicek M, Lenge o a M and
Sku ko a H (2021) Wo d
En opy-Based App oach o De ec
Highly Va iable Gene ic Ma ke s o
Bac e ial Geno yping.
F on . Mic obiol. 12:631605.
doi: 10.3389/ micb.2021.631605
Wo d En opy-Based App oach o
De ec Highly Va iable Gene ic
Ma ke s o Bac e ial Geno yping
Ma ke a Nyk yno a1*, Voj ech Ba on1, Ka el Sedla 1, Ma ej Bezdicek2,
Ma ina Lenge o a2and Helena Sku ko a1
1Depa men o Biomedical Enginee ing, Facul y o Elec ical Enginee ing and Communica ion, B no Uni e si y o Technology,
B no, Czechia, 2Depa men o In e nal Medicine - Hema ology and Oncology, Uni e si y Hospi al B no, B no, Czechia
Geno yping me hods a e used o dis inguish bac e ial s ains om one species. Thus,
dis inguishing bac e ial s ains on a global scale, be ween coun ies o local dis ic s
in one coun y is possible. Howe e , he highly selec ed bac e ial popula ions (e.g.,
local popula ions in hospi al) a e ypically closely ela ed and low di e si ied. The e o e,
cu en ly used yping me hods a e no able o dis inguish indi idual s ains om each
o he . He e, we p esen a no el pipeline o de ec highly a iable gene ic segmen s o
geno yping a closely ela ed bac e ial popula ion. The me hod is based on a deg ee o
diso de in analyzed sequences ha can be ep esen ed by sequence en opy. Wi h he
iden i ied a iable sequences, i is possible o ind ou ansmission ou es and sou ces
o highly i ulen and mul i esis an s ains. The p oposed me hod can be used o
any bac e ial popula ion, and due o i s whole genome ange, also non-coding egions
a e examined.
Keywo ds: geno yping, en opy, gene ic ma ke s, closely ela ed bac e ia, MLST
1. INTRODUCTION
Heal hca e-associa ed in ec ions (HAIs) a e a se ious wo ldwide h ea wi h signi ican impac on
mo ali y and mo bidi y o hospi alized pa ien s (A e ian e al., 2016). They a e de ined as in ec ions
de eloped in a hospi al o o he heal h ca e acili y ha i s appea 48 h o mo e a e hospi al
admission, o wi hin 30 days a e ha ing ecei ed heal h ca e (Haque e al., 2018). The e o e, he
abili y o iden i y he sou ce o he in ec ion and moni o he sp ead o disease is essen ial. To each
ha goal, a p ocess called bac e ial yping is used (Ruppi sch, 2016) as i can dis inguish indi idual
bac e ial s ains om one species. Howe e , apid labo a o y and compu a ional echniques’
esolu ion o yping is limi ed and he abili y o dis inguish s ains among a low di e si y hospi al
bac e ial popula ion is p ac ically impossible. Me hods such as pulsed- ield gel elec opho esis
(PFGE) (Schwa z and Can o , 1984), epe i i e sequence-based PCR ( ep-PCR) (Sku ko a e al.,
2019), mul ilocus sequence yping (MLST) (Maiden e al., 1998), o mini-MLST (Bezdicek e al.,
2019) a e used o cha ac e ize and in es iga e ou b eaks. As housekeeping genes used in MLST
o mini-MLST may no ha e su icien a iabili y in some cases, he e is s ill a need o de elop
new app oaches o iden i y a iable sequences. One o hem is mul ispace yping (MST) which
uses sequences occu ing in junk DNA (in e genic egions and pseudogenes) as hey show highe
a iabili y han coding genome pa s (Foucaul e al., 2005). Ano he app oach is single-locus
sequence yping (SLST), whe e one a iable sequence iden i ied based on a genome mining
app oach can ha e simila disc imina o y powe o MLST (Scholz e al., 2014). Howe e , i p ecise
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
analyses o he low di e si ied bac e ial popula ion a e equi ed,
he only me hod ha can be used is whole-genome sequencing
(WGS). Fo WGS yping, wo main app oaches exis : single
nucleo ide a ian analysis (SNV) and co e genome MLST
(cgMLST) o whole genome MLST (wgMLST) (Hen i e al., 2017;
Schü ch e al., 2018), whe e genes common o bac e ial isola es
a e compa ed. Ne e heless, he yping me hods which equi e
WGS a e no sui able o ou ine clinical p ac ice. Addi ionally,
he compa ison o di e en s udies is di icul as di e en da a
quali y assessmen , genome assembly, and esul s analysis a e
done (Ruppi sch, 2016).
In his a icle, we p esen a new app oach o localize highly
a iable sequences in bac e ial genomes o Esche ichia coli
and En e ococcus aecium. De ec ing new yping agmen s is
based on a calcula ion o wo d en opy in genomes assembled
om NGS da a. Po en ial a iable sequences a e e alua ed by
phylogene ic analysis, and he ob ained esul s a e compa ed o
cgMLST. The co e genome was chosen as he iden i ied ma ke s
mus be p esen ed in all bac e ial s ains o gi en species. Fo
ha eason, i is no app op ia e o s udy he pan-genome as i
con ains he co e genome and he dispensable genome (Medini
e al., 2005). The iden i ied ma ke s will be used o closely ela ed
bac e ial yping ia s anda d labo a o y me hods wi hou he
need o u he WGS. Ou app oach can be used o any bac e ia;
he only equi emen is aligned isola e genomes.
2. MATERIALS AND METHODS
2.1. Da ase
Fo his a icle, samples we e collec ed in he Uni e si y
Hospi al o B no. In o al, 23 E. aecium isola es we e ob ained
om se en hospi al depa men s be ween 6/2017 and 2/2018
and 21 E. coli isola es we e collec ed om a single hospi al
depa men be ween 5/2019 and 7/2019. The e o e, he da ase s
ha e di e en a iabili y a es as hey we e collec ed du ing
di e en pe iods and om a di e en numbe o depa men s.
Sequencing lib a ies we e p epa ed wi h KAPA Hype Plus
Ki s (Roche, Swi ze land), and a quali y check was pe o med
using a 2100 Bioanalyze (Agilen Technologies, USA). Lib a ies
we e quan i ied wi h KAPA Lib a y Quan i ica ion ki (Roche,
Swi ze land). Sequencing was pe o med on an Illumina MiSeq
pla o m using MiSeq Reagen Ki 2 (500-cycles) (Illumina,
USA) and pai ed-end eads (250 bp) we e acqui ed.
2.2. Re e ence-Based Genome Mapping
A e he sequenced da a’s quali y con ol, he eads we e
mapped ia BBMap ( 38.71, Bushnell, 2014) o he human
genome (GRCh38.p13) o emo e con amina ion which could
eme ge du ing sample p epa a ion. Reads which did no map
o he human genome we e u he analyzed and quali y
assessmen was done. Adap e s and low-quali y pa s o eads
we e immed ia T immoma ic ( 0.36, Bolge e al., 2014).
The eads leng h was abou 250 bp; hus he e e ence-based
genome mapping app oach was chosen, and o his pu pose, he
Bu ows-Wheele Alignmen MEM ( 0.7.17- 1188, Li, 2013) was
used. The e e ence sequences we e ob ained om he GenBank
da abase (Cla k e al., 2016). As e e ence sequence o assembly
o E. aecium genomes was chosen CP003351.1 (Lam e al.,
2012) and E. coli genomes we e assembled agains he sequence
BA000007.3 (Hayashi, 2001). Then Sam ools ( 1.7, Li e al.,
2009) was used o emo e low-quali y and duplica ed eads, and
consensus sequences we e gene a ed.
2.3. The En opy-Based De ec ion o
Va iable Sequences
To loca e highly a iable egions in sequences, he en opy alue
can be used. The DNA sequences en opy es ima ion (Schmi
and He zel, 1997) is based on he Shannon-en opy (Shannon,
1948). In his pape , he en opy calcula ion modi ica ion as
in e se alue o in o ma ion con en is u ilized. This modi ica ion
commonly used o sequence logo de e mina ion (Schneide and
S ephens, 1990) measu es he conse a ion o posi ion in bi s
which be e espec s di e en cha se s wi h a iable con en o
ambi alen bases.
En opy Hi(Schneide and S ephens, 1990) can be
calcula ed as
Hi= − X
a=A,C,G,T
a,i·log2 a,i, (1)
whe e iis he posi ion in he sequence, is he equency o one
o ou nucleo ides’ occu ence (a= A, C, G, T) in posi ion iand
can be calcula ed as
a,i=na
n, (2)
whe e nais numbe o occu ences o one nucleo ide in one
posi ion and nis numbe o all cha ac e s a he gi en posi ion.
In he ideal case, we wan o ind such a egion capable o
dis inguishing each isola e. Fo ha eason, he cu en me hod
o en opy calcula ion in one posi ion (Nyk yno a e al., 2019)
in he sequence’s alignmen canno be used, as i can only
dis inguish ou a ian s because ou nucleo ides (A, C, G, T)
a e p esen ed. In his a icle, a new unique app oach using wo d
en opy is p esen ed, whe e he wo d is a nucleo ide sequence o
de ined leng h.
As e e ence-based genome assembly is no pe ec du ing
consensus calling ambiguous nucleo ides can appea in he
consensus sequences, and hey a i icially inc ease he en opy
alue as shown in Figu e 1. To p e en de ec ing alsely a iable
sequences, he whole columns o aligned sequences whe e
ambiguous nucleo ides occu we e emo ed and o p ese e he
o iginal leng h o he c ea ed consensus sequences eplaced by
gaps. Also, he de ec ed ma ke s should occu in all analyzed
isola es. Thus, when gaps appea in some sequences, he gaps
ep esen ed by dashes a e added o he whole columns as is also
shown in Figu e 1.
In he window w, whose size co esponds o wo d leng h, he
numbe o unique wo ds is de e mined, and hen he en opy o
wo ds Hwis calcula ed as
Hw= − X(kw
n·(log2
kw
n)), (3)
F on ie s in Mic obiology | www. on ie sin.o g 2Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
FIGURE 1 | Example o sequence p ep ocessing. (A) Example o aligned sequences be o e p ep ocessing and hei en opy o wo d alues o wo ds o leng h i e.
(B) Aligned sequences a e p ep ocessing and co esponding en opy o wo d alues.
FIGURE 2 | Scheme o en opy o wo ds’ calcula ion o a se o sequences a e p ep ocessing and o wo d leng h o 5 bp and one nucleo ide window shi .
whe e kwis he numbe o occu ences o unique wo ds in each
alignmen in he cu en window, and nis he numbe o isola es.
The sum is applied o e he alignmen in he cu en window.
A e es ima ing he en opy alue, he window mo es wi h a shi
o one nucleo ide. The p inciple o en opy o wo ds’ calcula ion
is shown in Figu e 2.
Fo de ined sliding window sizes w(also called wo d leng h),
he en opy signal wi h leng h N−w+1 is ob ained, whe e Nis
he analyzed genome’s leng h.
In he signal, he maximum alue is iden i ied. In heo y,
he maximum alue can be de ined as a bina y loga i hm
o he numbe o sequences. Based on he maximum alue
F on ie s in Mic obiology | www. on ie sin.o g 3Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
ha occu s in he signal, he h eshold is se as a maximum
alue educed by 0.4 bi , and signal peaks abo e he h eshold
a e labeled as po en ial gene ic ma ke s. The s a ed alue
was chosen empi ically based on he numbe o po en ial
a iable egions o ensu e ha some egions will ha e su icien
disc imina o y powe ; howe e , i can be adjus ed o analyzed
da a by he use . The aligned sequences co esponding o
peaks a e ex ac ed and analyzed. I wo peaks a e dis an
less han he window leng h, he sequences co esponding
o he peaks a e me ged and analyzed as one. De eloped
sou ce codes and unc ions in Ma lab R2020a can be ound
a h ps://gi hub.com/ma ke anyk yno a/ma ke s_de ec ion.
2.4. Phylogene ic Analysis
In he i s pa , each se o aligned sequences co esponding o
peaks was analyzed. The e olu iona y dis ances we e calcula ed
based on he Kimu a (Kimu a, 1980) model, which is used as
s anda d. F om ob ained dis ances, he phylogene ic ees we e
cons uc ed based on UPGMA. Clus e analysis was conduc ed,
and he clus e s om he ees we e es ablished and compa ed
o esul s ob ained om cgMLST. The h eshold o clus e s
de e mina ion was se o in e sec he ee so ha he maximum
numbe o clus e s wasob ained.
The esul s showed ha no se had su icien disc imina o y
powe o dis inguish be ween s ains. Tha is why he
combina ion o wo aligned sequences se s was used. The
iden i ied a iable se s o sequences we e me ged, which c ea ed
one long sequence o each analyzed bac e ial isola e. Then
o me ged sequences, again he e olu iona y dis ances we e
calcula ed and based on hem, he phylogene ic ees we e
cons uc ed and examined. The leng h o a a iable se o
sequences had a p opo ional impac on he c ea ed phylogene ic
ee. I he sequences in one a iable se had a longe leng h,
hey had a highe impac on he ee and ice e sa, he se
o sho e sequences did no a ec he inal ee so much.
Ne e heless, sequences a iabili y has he main impac on he
phylogene ic ee.
3. RESULTS AND DISCUSSION
3.1. Sequence Type De e mina ion and
cgMLST Analysis
The inco po a ed E. coli and E. aecium cgMLST, cgSNV, and
wgSNV schemes in Ridom SeqSphe e+ so wa e (Ridom, DE)
we e used o analyze he gene ic ela edness o sequenced
bac e ial genomes. Fo E. coli, his included 3,152 a ge genes
used o gene a e a cgMLST dend og am. Fo he cgSNV
minimum spanning ee analysis, 152,212 aligned nucleo ide si es
we e analyzed. In o al, 176,629 aligned nucleo ide si es we e used
o wgSNV minimum spanning ee analysis. Fo E. aecium, he
scheme included 1,423 a ge genes o cgMLST, 7,707 aligned
nucleo ide si es o cgSNV and 10,858 aligned nucleo ide si es
o wgSNV analysis. The WGS da a we e used o de e mine
sequence ypes (STs) o bo h, E. coli and E. aecium s ains. Fo E.
coli, se en housekeeping genes (adk, umC, gy B, lcd, mdh, pu A,
ecA) and o E. aecium again se en housekeeping genes (a pA,
ddl, gdh, pu K, gyd, ps S, adk) we e analyzed. Fo E. aecium
h ee STs we e iden i ied (4 ×ST 117, 1 ×ST 18, 18 ×ST 80),
and o E. coli ele en STs we e ecognized and one ST emains
unknown (1 ×ST 69, 4 ×ST 131, 1 ×ST 95, 2 ×ST 404, 2
×ST 38, 2 ×ST 1049, 4 ×ST 58, 1 ×ST 297, 1 ×ST 517, 2
×ST 101, 1 ×ST UNW). The comple e esul s a e a ached as
Supplemen a y File 1.
3.2. Po en ial Gene ic Ma ke s Analysis
Sequences wi h a high alue o wo d en opy we e iden i ied
o wo d leng hs om 50 o 400 bp wi h a s ep o 50 bp.
T adi ional labo a o y me hods can use iden i ied ma ke s o
leng h om his ange. In he i s pa , one a iable agmen
wi h high wo d en opy alues was used o dis inguish he
s ains o E. coli and E. aecium. Examining 1,358 agmen s
wi h high a iabili y o all wo d sizes showed ha he
maximum numbe o clus e s ha can be dis inguished o
E. coli was 13. Fo E. aecium, 104 po en ial gene ic ma ke s
o all wo d sizes we e analyzed, and he maximum numbe
o dis inguishable clus e s was 5. The comple e esul s a e in
Supplemen a y File 2.
The numbe o dis inguishable clus e s using one a iable
agmen is no high enough. Fo ha eason, wo a iable
agmen s wi h wo d en opy abo e h eshold we e analyzed.
Fo all combina ions o wo possible gene ic ma ke s, he
phylogene ic ees we e cons uc ed, and he numbe o clus e s
was calcula ed. In he nex s ep, he po en ial gene ic ma ke s
wi h he highes numbe o clus e s we e analyzed. Fo all wo d
sizes, 136,661 combina ions o possible gene ic ma ke s o E. coli
and 685 o E. aecium we e examined in o al.
In case o need, he combina ion o e en mo e ma ke s
can be used; howe e , he ime o analysis will g ow as
he ime complexi y is O(nm), whe e mis he numbe o
a iable agmen s.
Fo E. coli, 29 ees we e examined which di ided he isola es
o 15 clus e s. Ten ees classi y he genomes acco ding o he
11 p ede ined clus e s om cgMLST. The esul s om cgMLST
wi h 11 ma ked sequence ypes and 11 de ined clus e s and a
cladog am c ea ed based on he combina ion o wo gene ic
ma ke s wi h 15 labeled clus e s a e shown in Figu e 3.
In each case o one gene ic ma ke , he egion which
con ained he genes o hypo he ical p o ein (NP_308436.2) and
S- o mylglu a hione hyd olase (NP_308437.1) was selec ed. In
i e ou o en cases, he in e genic egion in on o hypo he ical
p o ein was also inco po a ed. As he second gene ic ma ke ,
wo egions we e mos o en ( h ee imes) chosen. The i s one
con ained ou e memb ane p o ein (NP_313225.1) and in one
case he in e genic egion in on o he gene oo. The second
one s a ed in imb ial-like adhesin p o ein (NP_310944.1) and
ended in hypo he ical p o ein (NP_310945.1). Fou o he gene ic
ma ke s we e selec ed only one ime. Two o hem consis ed
o non-coding egion and gene [PhoH amily P-loop ATPase
(NP_309293.1) in he i s case, excinuclease U ABC subuni
U C (NP_310678.2) in he second case], o ano he ma ke
pa o he gene alyl- RNA syn he ase (NP_313262.1) was
picked, and in he las case, as a gene ic ma ke he non-
coding egion was selec ed. De ailed in o ma ion is a ailable in
F on ie s in Mic obiology | www. on ie sin.o g 4Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
FIGURE 3 | Classi ica ion o 21 E. coli genomes. (A) Co e genome MLST analysis o 21 E. coli isola es wi h p ede ined ma ked clus e s (ou e ci cles) and sequence
ypes (inne ci cles). (B) Cladog am o 21 E. coli isola es based on wo iden i ied gene ic ma ke s wi h highligh clus e s ob ained om co e genome MLST analysis
(ou e ci cle) and subg oups ob ained based on analysis o iden i ied a iable sequences (inne ci cle) c ea ed by E ol iew (Sub amanian e al., 2019).
FIGURE 4 | Classi ica ion o 23 E. aecium genomes. (A) Co e genome MLST analysis o 23 E. aecium isola es wi h p ede ined ma ked clus e s (ou e ci cles) and
sequence ypes (inne ci cles). (B) Cladog am o 23 E. aecium isola es based on wo iden i ied gene ic ma ke s wi h highligh clus e s ob ained om co e genome
MLST analysis (ou e ci cle) and subg oups ob ained based on analysis o iden i ied a iable sequences (inne ci cle) c ea ed by E ol iew (Sub amanian e al., 2019).
Supplemen a y File 2 and all phylogene ic ees ha co ec ly
classi ied genomes a e depic ed in Supplemen a y Figu e 1.
Fo E. aecium, 22 ees which di ided he isola es in o nine
clus e s we e examined. I was ound ha six ees classi y he
genomes acco ding o i e p ede ined clus e s om cgMLST. In
Figu e 4, he esul s om cgMLST analysis wi h h ee labeled
sequence ypes, i e clus e s and a cladog am cons uc ed based
on he iden i ied combina ion o wo gene ic ma ke s wi h nine
labeled clus e s shown.
As in he p e ious case, in all possible combina ions,
one gene ic ma ke always emained he same, and i was
RNA me hyl ans e ase (WP_002294889.1). In one o six cases,
he in e genic egion a ound he gene was also inco po a ed
in o he a iable gene ic pa . Mg- ansloca ing P- ype ATPase
(WP_002304253.1) was in mos cases ( h ee imes) selec ed as
he second ma ke . The egion ha con ained he gene o
DUF1430 domain-con aining p o ein (WP_002288541.1) was
chosen wice. Fo he las possible gene ic ma ke , he gene o
nucleoside-diphospha e kinase (WP_002293327.1) was picked.
The comple e esul s a e a ailable in Supplemen a y File 2 and
all phylogene ic ees co ec ly classi ying genomes a e shown in
Supplemen a y Figu e 2.
3.3. Wo d Sizes
Fo de ined wo d sizes, a maximum numbe o co ec ly classi ied
clus e s o phylogene ic ees based on wo ma ke combina ions
F on ie s in Mic obiology | www. on ie sin.o g 5Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
FIGURE 5 | Analysis o wo d size impac on numbe o co ec ly classi ied clus e s. (A) Maximum numbe o co ec ly classi ied clus e s o combina ions o wo
a iable agmen s o di e en wo d sizes and numbe o wo a iable agmen combina ions ha co ec ed classi ied analyzed E. aecium isola es. (B) Maximum
numbe o co ec ly classi ied clus e s o combina ion o wo a iable agmen s o di e en wo d sizes and numbe o wo a iable agmen combina ions ha
co ec ed classi ied analyzed E. coli isola es.
was es ablished, as is shown in Figu e 5. Fo analyzed ees he
numbe o clus e s was de e mined and he co ec genomes
classi ica ion was con olled. The esul s showed ha o E.
aecium he wo d leng hs do no signi ican ly a ec he maximum
numbe o co ec ly es ablished clus e s, which is nine o all
wo d sizes excep he leng h 50 and 300 bp. On he con a y, he
esolu ion powe o E. coli g ows wi h inc easing wo d leng hs
and he bes esul s (dis inguish 15 clus e s) can be ob ained o a
window o size 300 bp and mo e.
3.4. Resul s Ve i ica ion
The analyzed da ase s we e ex ended by comple e genomes
sequences o E. coli and E. aecium ob ained om he
GenBank da abase. Fo each bac e ium, six sequences we e
downloaded, and he STs we e de e mined. Fo E. coli
isola es (CP050212.1, CP050219.1, CP050211.1, CP050201.1,
CP050214.1, CP050205.1) i e STs we e iden i ied (1 ×ST 2705,
1×ST 2280, 1 ×ST 70, 1 ×ST 602, 2 ×ST 131), and
o E. aecium isola es (CP040236.1, CP036151.1, CP027501.1,
CP035666.1, CP035660.1, CP035220.1) one ST was ecognized,
and one ST emains unknown (5 ×ST 80, 1 ×ST UNW).
The esul s a e in Supplemen a y File 1. Using he BLAST+
( 2.6.0+, Camacho e al., 2009) he iden i ied a iable ma ke s
we e loca ed in hese genomes. The ma ke s sequences o
bo h se s (o iginal and da abase) we e analyzed oge he . The
sequences we e aligned, e olu iona y dis ances we e calcula ed,
and phylogene ic ees we e cons uc ed. To ally, 27 isola es o
E. coli belonging o nine STs we e classi ied o 17 clus e s and
29 isola es o E. aecium belonging o ou STs we e i e imes
classi ied o 11 clus e s and once o 10 clus e s. All c ea ed
phylogene ic ees a e depic ed in Supplemen a y Figu e 3.
The iden i ied gene ic ma ke s can be used o yping no el
isola es. Howe e , i a la ge numbe o new isola es should
be analyzed, he algo i hm should be un again o ind
ma ke s wi h he highes disc imina o y powe o he analyzed
bac e ial popula ion.
3.5. Analyzed Bac e ia
The p oposed app oach was es ed on wo bac e ia—E. aecium
and E. coli which belong among signi ican HAIs pa hogens.
E. aecium is a G am-posi i e bac e ium and can be a sou ce
o nosocomial in ec ions, especially in immunocomp omised
pa ien s, and causes endoca di is, bac e emia o u ina y ac
in ec ion. In addi ion, some E. aecium isola es can de elop d ug
esis ance o se e al an ibio ics g oups such as glycopep ides,
be a-lac ams, luo oquinolones, o aminoglycosides (Yoong
e al., 2004; Cas illo-Rojas e al., 2013). Ano he pa hogen
om G am-nega i e bac e ia is E. coli, which can be ound
in human in es inal lo a as a commensal. Howe e , as an
oppo unis ic pa hogen E. coli ep esen s a huge public heal h
p oblem. I is one o he mos common causes o HAIs
ypically connec ed o meningi is, pneumonia, u ina y ac
in ec ions, and so issue in ec ions (Jau eguy e al., 2008). E.
coli s ains ha e he abili y o accumula e esis ance genes; hus,
he esis ance o b oad-spec um cephalospo ins, ca bapenems,
aminoglycoside, luo oquinolones, and polymyxins is o en
obse ed (Poi el e al., 2018).
F on ie s in Mic obiology | www. on ie sin.o g 6Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
4. CONCLUSION
As yping is c ucial o iden i ying he sou ce o in ec ion
and ou b eak moni o ing, i is necessa y o ha e a highly
sensi i e yping me hod o dis inguish closely ela ed bac e ial
s ains. Howe e , cu en ly used me hods ha e insu icien
disc imina o y powe excep whole genome sequencing.
Ne e heless, WGS is ime-consuming, and da a p ocessing is
no possible in ou ine clinical p ac ice. Fo his eason, a new
app oach o de ec a iable gene ic ma ke s om WGS da a is
p esen ed. The WGS is only used once o ob ain he i s se
o da a o gene ic ma ke s de ec ion. To loca e highly a iable
sequences, wo d en opy is employed. The loca ed ma ke s can
be used o yping a closely ela ed bac e ial popula ion ia
s anda d labo a o y me hods. Thus, ano he genome sequencing
will be necessa y only i a la ge numbe o new isola es should
be analyzed.
Using he en opy-based app oach, new gene ic ma ke s wi h
he same o e en highe disc imina o y powe han cgMLST can
be iden i ied. As he leng h o wo ds om 50 o 400 bp wi h a
s ep o 50 bp is applied, po en ial ma ke s can be used in s anda d
labo a o y echniques because hey ha e sui able leng h. The only
equi emen o he p oposed app oach is ha genomes mus be
aligned, which can cause a p oblem mainly when la ge genomes
should be examined. The p esen ed me hod was es ed on bo h
G am-posi i e and G am-nega i e bac e ia.
As in bo h es ed da ase s, one ma ke wi h su icien
disc imina o y powe o yping o low di e si y bac e ial
popula ion was no ound, he combina ion o wo ma ke s
was employed. Based on wo iden i ied ma ke s, we we e able
o dis inguish he G am-nega i e and G am-posi i e bac e ial
popula ion a he same and e en a highe disc imina ion le el
han ia classic labo a o y me hods using only wo a iable
sequences ins ead o se en housekeeping genes. The iden i ied
ma ke s will be used in ou ine labo a o y me hods o bac e ial
yping. The low numbe o a iable agmen s ha should be
examined will sho en he ime o analysis.
Commonly used yping me hods mainly use coding egions
and analyze genes conse ed and p esen ed in all isola es
o he gi en species. Hence he a iabili y a e is lowe and
yping he closely ela ed popula ion is no possible. The
p oposed me hod uses he whole genome ange; hus, also
in e genic genome egions a e analyzed as hey ha e highe
a iabili y han coding a eas and can con ain gene ic ma ke s
o yping ela ed isola es. Based on iden i ied ma ke s, he
closely ela ed bac e ial popula ion can be di e si ied and o
ind ou he ansmission ou es and sou ce o ou b eaks will
be possible.
DATA AVAILABILITY STATEMENT
The da ase s p esen ed in his s udy can be ound in online
eposi o ies. The names o he eposi o y/ eposi o ies and
accession numbe (s) can be ound a : h ps://www.ncbi.nlm.nih.
go /, PRJNA675431.
AUTHOR CONTRIBUTIONS
MN, VB, KS, and HS con ibu ed o he concep ion and design
o he s udy. MN implemen ed he algo i hm and e alua ed he
esul s. ML and MB ensu ed he biological aspec s o he p ojec .
All au ho s ead and app o ed he inal manusc ip .
FUNDING
This wo k was suppo ed by a g an p ojec om he Czech
Science Founda ion (GA17-01821S).
SUPPLEMENTARY MATERIAL
The Supplemen a y Ma e ial o his a icle can be ound
online a : h ps://www. on ie sin.o g/a icles/10.3389/ micb.
2021.631605/ ull#supplemen a y-ma e ial
Supplemen a y File 1 | MLST analysis esul s o E.coli and E. aecium.
Supplemen a y File 2 | Po en ial and inal gene ic ma ke s o E. coli
and E. aecium.
Supplemen a y Figu e 1 | Phylogene ic ees ha co ec ly classi ied E. coli
genomes wi h ma ked clus e s.
Supplemen a y Figu e 2 | Phylogene ic ees ha co ec ly classi ied E. aecium
genomes wi h ma ked clus e s.
Supplemen a y Figu e 3 | Phylogene ic ees o o iginal and da abase genomes
o E. aecium and E. coli genomes wi h ma ked clus e s.
REFERENCES
A e ian, H., Vogel, M., Kwe ka , A., and Ha mann, M. (2016). Economic
e alua ion o in e en ions o p e en ion o hospi al acqui ed in ec ions: a
sys ema ic e iew. PLoS ONE 11:e0146381. doi: 10.1371/jou nal.pone.0146381
Bezdicek, M., Nyk yno a, M., Ple o a, K., B helo a, E., Kocmano a, I.,
Sedla , K., e al. (2019). Applica ion o mini-MLST and whole genome
sequencing in low di e si y hospi al ex ended-spec um be a-lac amase
p oducing Klebsiella pneumoniae popula ion. PLoS ONE 14:e0221187.
doi: 10.1371/jou nal.pone.0221187
Bolge , A. M., Lohse, M., and Usadel, B. (2014). T immoma ic: a lexible
imme o Illumina sequence da a. Bioin o ma ics 30, 2114–2120.
doi: 10.1093/bioin o ma ics/b u170
Bushnell, B. (2014). BBMap: A Fas , Accu a e, Splice-Awa e Aligne . No. LBNL-
7065E. Be keley, CA: E nes O lando Law ence Be keley Na ional Labo a o y.
Camacho, C., Coulou is, G., A agyan, V., Ma, N., Papadopoulos, J., Beale , K., e al.
(2009). BLAST+: a chi ec u e and applica ions. BMC Bioin o ma ics 10:421.
doi: 10.1186/1471-2105-10-421
Cas illo-Rojas, G., Maza i-Hi ía , M., Ponce de León, S., Amie a-Fe nández,
R. I., Agis-Juá ez, R. A., Huebne , J., e al. (2013). Compa ison
o En e ococcus aecium and En e ococcus aecalis s ains isola ed
om wa e and clinical samples: an imic obial suscep ibili y and
gene ic ela ionships. PLoS ONE 8:e59491. doi: 10.1371/jou nal.pone.
0059491
Cla k, K., Ka sch-Miz achi, I., Lipman, D. J., Os ell, J., and Saye s, E. W. (2016).
GenBank. Nucleic Acids Res. 44, D67–D72. doi: 10.1093/na /gk 1276
Foucaul , C., La Scola, B., Lind oos, H., Ande sson, S. G. E., and Raoul ,
D. (2005). Mul ispace yping echnique o sequence-based yping o
Ba onella quin ana. J. Clin. Mic obiol. 43, 41–48. doi: 10.1128/JCM.43.1.41-4
8.2005
F on ie s in Mic obiology | www. on ie sin.o g 7Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
Haque, M., Sa elli, M., McKimm, J., and Abu Baka , M. B. (2018). Heal h
ca e-associa ed in ec ions-an o e iew. In ec . D ug Resis . 11, 2321–2333.
doi: 10.2147/IDR.S177247
Hayashi, T. (2001). Comple e genome sequence o en e ohemo hagic Eschelichia
coli O157:H7 and genomic compa ison wi h a labo a o y s ain K-12. DNA Res.
8, 11–22. doi: 10.1093/dna es/8.1.11
Hen i, C., Leeki cha oenphon, P., Ca le on, H. A., Radomski, N., Kaas, R. S.,
Ma ie , J.-F., e al. (2017). An assessmen o di e en genomic app oaches
o in e ing phylogeny o Lis e ia monocy ogenes. F on . Mic obiol. 8:2351.
doi: 10.3389/ micb.2017.02351
Jau eguy, F., Land aud, L., Passe , V., Diancou , L., F apy, E., Guigon, G., e al.
(2008). Phylogene ic and genomic di e si y o human bac e emic Esche ichia
coli s ains. BMC Genomics 9:560. doi: 10.1186/1471-2164-9-560
Kimu a, M. (1980). A simple me hod o es ima ing e olu iona y a es o base
subs i u ions h ough compa a i e s udies o nucleo ide sequences. J. Mol. E ol.
16, 111–120. doi: 10.1007/BF01731581
Lam, M. M. C., Seemann, T., Bulach, D. M., Gladman, S. L., Chen, H., Ha ing, V.,
e al. (2012). Compa a i e analysis o he i s comple e En e ococcus aecium
genome. J. Bac e iol. 194, 2334–2341. doi: 10.1128/JB.00259-12
Li, H. (2013). Aligning sequence eads, clone sequences and assembly con igs wi h
BWA-MEM. a Xi [P ep in ]. a Xi :1303.3997
Li, H., Handsake , B., Wysoke , A., Fennell, T., Ruan, J., Home , N., e al.
(2009). The Sequence Alignmen /Map o ma and SAM ools. Bioin o ma ics 25,
2078–2079. doi: 10.1093/bioin o ma ics/b p352
Maiden, M. C. J., Byg a es, J. A., Feil, E., Mo elli, G., Russell, J. E., U win, R., e al.
(1998). Mul ilocus sequence yping: a po able app oach o he iden i ica ion o
clones wi hin popula ions o pa hogenic mic oo ganisms. P oc. Na l. Acad. Sci.
U.S.A. 95, 3140–3145. doi: 10.1073/pnas.95.6.3140
Medini, D., Dona i, C., Te elin, H., Masignani, V., and Rappuoli, R.
(2005). The mic obial pan-genome. Cu . Opin. Gene . De . 15, 589–594.
doi: 10.1016/j.gde.2005.09.006
Nyk yno a, M., Made anko a, D., Ba on, V., Bezdicek, M., Lenge o a, M., and
Sku ko a, H. (2019). “En opy-based de ec ion o gene ic ma ke s o bac e ia
geno yping,” in Lec u e No es in Compu e Science (including subse ies Lec u e
No es in A i icial In elligence and Lec u e No es in Bioin o ma ics) (Cham:
Sp inge ), 11466 LNBI, 177–188. doi: 10.1007/978-3-030-17935-9_17
Poi el, L., Madec, J.-Y., Lupo, A., Schink, A.-K., Kie e , N., No dmann, P.,
e al. (2018). “An imic obial esis ance in Esche ichia coli,” in An imic obial
Resis ance in Bac e ia om Li es ock and Companion Animals, eds S.
Schwa z, L. M. Ca aco, and J. Shen (Washing on, DC: Ame ican Socie y o
Mic obiology), 289–316. doi: 10.1128/mic obiolspec.ARBA-0026-2017
Ruppi sch, W. (2016). Molecula yping o bac e ia o epidemiological
su eillance and ou b eak in es iga ion/Molekula e Typisie ung on
Bak e ien ü die epidemiologische Übe wachung und Ausb uchsabklä ung.
J. Land Manage. Food En i on. 67, 199–224. doi: 10.1515/boku-20
16-0017
Schmi , A. O., and He zel, H. (1997). Es ima ing he en opy o DNA sequences.
J. Theo . Biol. 188, 369–377. doi: 10.1006/j bi.1997.0493
Schneide , T. D., and S ephens, R. M. (1990). Sequence logos: a new
way o display consensus sequences. Nucl. Acids Res. 18, 6097–6100.
doi: 10.1093/na /18.20.6097
Scholz, C. F. P., Jensen, A., Lomhol , H. B., B üggemann, H., and Kilian, M.
(2014). A no el high- esolu ion single locus sequence yping scheme o
mixed popula ions o P opionibac e ium acnes in i o. PLoS ONE 9:e104199.
doi: 10.1371/jou nal.pone.0104199
Schü ch, A., A edondo-Alonso, S., Willems, R., and Goe ing, R. (2018). Whole
genome sequencing op ions o bac e ial s ain yping and epidemiologic
analysis based on single nucleo ide polymo phism e sus gene-by-gene-based
app oaches. Clin. Mic obiol. In ec . 24, 350–354. doi: 10.1016/j.cmi.2017.12.016
Schwa z, D. C., and Can o , C. R. (1984). Sepa a ion o yeas ch omosome-
sized DNAs by pulsed ield g adien gel elec opho esis. Cell 37, 67–75.
doi: 10.1016/0092-8674(84)90301-5
Shannon, C. E. (1948). A ma hema ical heo y o communica ion. Bell Sys . Techn.
J. 27, 379–423. doi: 10.1002/j.1538-7305.1948. b01338.x
Sku ko a, H., Vi ek, M., Bezdicek, M., B helo a, E., and Lenge o a, M. (2019).
Ad anced DNA inge p in geno yping based on a model de eloped om eal
chip elec opho esis da a. J. Ad . Res. 18, 9–18. doi: 10.1016/j.ja e.2019.01.005
Sub amanian, B., Gao, S., Le che , M. J., Hu, S., and Chen, W.-H. (2019). E ol iew
3: a webse e o isualiza ion, anno a ion, and managemen o phylogene ic
ees. Nucl. Acids Res. 47, W270-W275. doi: 10.1093/na /gkz357
Yoong, P., Schuch, R., Nelson, D., and Fische i, V. A. (2004). Iden i ica ion o a
b oadly ac i e phage ly ic enzyme wi h le hal ac i i y agains an ibio ic- esis an
En e ococcus aecalis and En e ococcus aecium. J. Bac e iol. 186, 4808–4812.
doi: 10.1128/JB.186.14.4808-4812.2004
Con lic o In e es : The au ho s decla e ha he esea ch was conduc ed in he
absence o any comme cial o inancial ela ionships ha could be cons ued as a
po en ial con lic o in e es .
Copy igh © 2021 Nyk yno a, Ba on, Sedla , Bezdicek, Lenge o a and Sku ko a.
This is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons
A ibu ion License (CC BY). The use, dis ibu ion o ep oduc ion in o he o ums
is pe mi ed, p o ided he o iginal au ho (s) and he copy igh owne (s) a e c edi ed
and ha he o iginal publica ion in his jou nal is ci ed, in acco dance wi h accep ed
academic p ac ice. No use, dis ibu ion o ep oduc ion is pe mi ed which does no
comply wi h hese e ms.
F on ie s in Mic obiology | www. on ie sin.o g 8Feb ua y 2021 | Volume 12 | A icle 631605