scieee Science in your language
[en] (orig)

Word entropy-based approach to detect highly variable genetic markers for bacterial genotyping

Abstract

Genotyping methods are used to distinguish bacterial strains from one species. Thus, distinguishing bacterial strains on a global scale, between countries or local districts in one country is possible. However, the highly selected bacterial populations (e.g. local populations in hospital) are typically closely related and low diversified. Therefore, currently used typing methods are not able to distinguish individual strains from each other. Here, we present a novel pipeline to detect highly variable genetic segments for genotyping a closely related bacterial population. The method is based on a degree of disorder in analyzed sequences that can be represented by sequence entropy. With the identified variable sequences, it is possible to find out transmission routes and sources of highly virulent and multiresistant strains. The proposed method can be used for any bacterial population, and due to its whole genome range, also noncoding regions are examined.

Read accessible full text

Word entropy-based approach to detect highly variable genetic markers for bacterial genotyping

Author: Jakubíčková, Markéta; Bartoň, Vojtěch; Sedlář, Karel; Bezdíček, Matěj; Lengerová, Martina; Vítková, Helena
Publisher: Frontiers Media SA
Year: 2021
DOI: 10.3389/fmicb.2021.631605
Source: https://dspace.vut.cz/bitstreams/c7b6207e-e1b1-454f-a8cd-e51348a239b4/download
ORIGINAL RESEARCH
published: 03 Feb ua y 2021
doi: 10.3389/ micb.2021.631605
F on ie s in Mic obiology | www. on ie sin.o g 1Feb ua y 2021 | Volume 12 | A icle 631605
Edi ed by:
Dominique La enie ,
UMR6074 Ins i u de Reche che en
In o ma ique e Sys èmes Aléa oi es
(IRISA), F ance
Re iewed by:
Mahend a Ma iadassou,
Ins i u Na ional de Reche che pou
l’Ag icul u e, l’Alimen a ion e
l’En i onnemen (INRAE), F ance
William She win,
Uni e si y o New Sou h Wales,
Aus alia
*Co espondence:
Ma ke a Nyk yno a
nyk yno a@ u b .cz
Special y sec ion:
This a icle was submi ed o
E olu iona y and Genomic
Mic obiology,
a sec ion o he jou nal
F on ie s in Mic obiology
Recei ed: 20 No embe 2020
Accep ed: 13 Janua y 2021
Published: 03 Feb ua y 2021
Ci a ion:
Nyk yno a M, Ba on V, Sedla K,
Bezdicek M, Lenge o a M and
Sku ko a H (2021) Wo d
En opy-Based App oach o De ec
Highly Va iable Gene ic Ma ke s o
Bac e ial Geno yping.
F on . Mic obiol. 12:631605.
doi: 10.3389/ micb.2021.631605
Wo d En opy-Based App oach o
De ec Highly Va iable Gene ic
Ma ke s o Bac e ial Geno yping
Ma ke a Nyk yno a1*, Voj ech Ba on1, Ka el Sedla 1, Ma ej Bezdicek2,
Ma ina Lenge o a2and Helena Sku ko a1
1Depa men o Biomedical Enginee ing, Facul y o Elec ical Enginee ing and Communica ion, B no Uni e si y o Technology,
B no, Czechia, 2Depa men o In e nal Medicine - Hema ology and Oncology, Uni e si y Hospi al B no, B no, Czechia
Geno yping me hods a e used o dis inguish bac e ial s ains om one species. Thus,
dis inguishing bac e ial s ains on a global scale, be ween coun ies o local dis ic s
in one coun y is possible. Howe e , he highly selec ed bac e ial popula ions (e.g.,
local popula ions in hospi al) a e ypically closely ela ed and low di e si ied. The e o e,
cu en ly used yping me hods a e no able o dis inguish indi idual s ains om each
o he . He e, we p esen a no el pipeline o de ec highly a iable gene ic segmen s o
geno yping a closely ela ed bac e ial popula ion. The me hod is based on a deg ee o
diso de in analyzed sequences ha can be ep esen ed by sequence en opy. Wi h he
iden i ied a iable sequences, i is possible o ind ou ansmission ou es and sou ces
o highly i ulen and mul i esis an s ains. The p oposed me hod can be used o
any bac e ial popula ion, and due o i s whole genome ange, also non-coding egions
a e examined.
Keywo ds: geno yping, en opy, gene ic ma ke s, closely ela ed bac e ia, MLST
1. INTRODUCTION
Heal hca e-associa ed in ec ions (HAIs) a e a se ious wo ldwide h ea wi h signi ican impac on
mo ali y and mo bidi y o hospi alized pa ien s (A e ian e al., 2016). They a e de ined as in ec ions
de eloped in a hospi al o o he heal h ca e acili y ha i s appea 48 h o mo e a e hospi al
admission, o wi hin 30 days a e ha ing ecei ed heal h ca e (Haque e al., 2018). The e o e, he
abili y o iden i y he sou ce o he in ec ion and moni o he sp ead o disease is essen ial. To each
ha goal, a p ocess called bac e ial yping is used (Ruppi sch, 2016) as i can dis inguish indi idual
bac e ial s ains om one species. Howe e , apid labo a o y and compu a ional echniques’
esolu ion o yping is limi ed and he abili y o dis inguish s ains among a low di e si y hospi al
bac e ial popula ion is p ac ically impossible. Me hods such as pulsed- ield gel elec opho esis
(PFGE) (Schwa z and Can o , 1984), epe i i e sequence-based PCR ( ep-PCR) (Sku ko a e al.,
2019), mul ilocus sequence yping (MLST) (Maiden e al., 1998), o mini-MLST (Bezdicek e al.,
2019) a e used o cha ac e ize and in es iga e ou b eaks. As housekeeping genes used in MLST
o mini-MLST may no ha e su icien a iabili y in some cases, he e is s ill a need o de elop
new app oaches o iden i y a iable sequences. One o hem is mul ispace yping (MST) which
uses sequences occu ing in junk DNA (in e genic egions and pseudogenes) as hey show highe
a iabili y han coding genome pa s (Foucaul e al., 2005). Ano he app oach is single-locus
sequence yping (SLST), whe e one a iable sequence iden i ied based on a genome mining
app oach can ha e simila disc imina o y powe o MLST (Scholz e al., 2014). Howe e , i p ecise
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
analyses o he low di e si ied bac e ial popula ion a e equi ed,
he only me hod ha can be used is whole-genome sequencing
(WGS). Fo WGS yping, wo main app oaches exis : single
nucleo ide a ian analysis (SNV) and co e genome MLST
(cgMLST) o whole genome MLST (wgMLST) (Hen i e al., 2017;
Schü ch e al., 2018), whe e genes common o bac e ial isola es
a e compa ed. Ne e heless, he yping me hods which equi e
WGS a e no sui able o ou ine clinical p ac ice. Addi ionally,
he compa ison o di e en s udies is di icul as di e en da a
quali y assessmen , genome assembly, and esul s analysis a e
done (Ruppi sch, 2016).
In his a icle, we p esen a new app oach o localize highly
a iable sequences in bac e ial genomes o Esche ichia coli
and En e ococcus aecium. De ec ing new yping agmen s is
based on a calcula ion o wo d en opy in genomes assembled
om NGS da a. Po en ial a iable sequences a e e alua ed by
phylogene ic analysis, and he ob ained esul s a e compa ed o
cgMLST. The co e genome was chosen as he iden i ied ma ke s
mus be p esen ed in all bac e ial s ains o gi en species. Fo
ha eason, i is no app op ia e o s udy he pan-genome as i
con ains he co e genome and he dispensable genome (Medini
e al., 2005). The iden i ied ma ke s will be used o closely ela ed
bac e ial yping ia s anda d labo a o y me hods wi hou he
need o u he WGS. Ou app oach can be used o any bac e ia;
he only equi emen is aligned isola e genomes.
2. MATERIALS AND METHODS
2.1. Da ase
Fo his a icle, samples we e collec ed in he Uni e si y
Hospi al o B no. In o al, 23 E. aecium isola es we e ob ained
om se en hospi al depa men s be ween 6/2017 and 2/2018
and 21 E. coli isola es we e collec ed om a single hospi al
depa men be ween 5/2019 and 7/2019. The e o e, he da ase s
ha e di e en a iabili y a es as hey we e collec ed du ing
di e en pe iods and om a di e en numbe o depa men s.
Sequencing lib a ies we e p epa ed wi h KAPA Hype Plus
Ki s (Roche, Swi ze land), and a quali y check was pe o med
using a 2100 Bioanalyze (Agilen Technologies, USA). Lib a ies
we e quan i ied wi h KAPA Lib a y Quan i ica ion ki (Roche,
Swi ze land). Sequencing was pe o med on an Illumina MiSeq
pla o m using MiSeq Reagen Ki 2 (500-cycles) (Illumina,
USA) and pai ed-end eads (250 bp) we e acqui ed.
2.2. Re e ence-Based Genome Mapping
A e he sequenced da a’s quali y con ol, he eads we e
mapped ia BBMap ( 38.71, Bushnell, 2014) o he human
genome (GRCh38.p13) o emo e con amina ion which could
eme ge du ing sample p epa a ion. Reads which did no map
o he human genome we e u he analyzed and quali y
assessmen was done. Adap e s and low-quali y pa s o eads
we e immed ia T immoma ic ( 0.36, Bolge e al., 2014).
The eads leng h was abou 250 bp; hus he e e ence-based
genome mapping app oach was chosen, and o his pu pose, he
Bu ows-Wheele Alignmen MEM ( 0.7.17- 1188, Li, 2013) was
used. The e e ence sequences we e ob ained om he GenBank
da abase (Cla k e al., 2016). As e e ence sequence o assembly
o E. aecium genomes was chosen CP003351.1 (Lam e al.,
2012) and E. coli genomes we e assembled agains he sequence
BA000007.3 (Hayashi, 2001). Then Sam ools ( 1.7, Li e al.,
2009) was used o emo e low-quali y and duplica ed eads, and
consensus sequences we e gene a ed.
2.3. The En opy-Based De ec ion o
Va iable Sequences
To loca e highly a iable egions in sequences, he en opy alue
can be used. The DNA sequences en opy es ima ion (Schmi
and He zel, 1997) is based on he Shannon-en opy (Shannon,
1948). In his pape , he en opy calcula ion modi ica ion as
in e se alue o in o ma ion con en is u ilized. This modi ica ion
commonly used o sequence logo de e mina ion (Schneide and
S ephens, 1990) measu es he conse a ion o posi ion in bi s
which be e espec s di e en cha se s wi h a iable con en o
ambi alen bases.
En opy Hi(Schneide and S ephens, 1990) can be
calcula ed as
Hi= − X
a=A,C,G,T
a,i·log2 a,i, (1)
whe e iis he posi ion in he sequence, is he equency o one
o ou nucleo ides’ occu ence (a= A, C, G, T) in posi ion iand
can be calcula ed as
a,i=na
n, (2)
whe e nais numbe o occu ences o one nucleo ide in one
posi ion and nis numbe o all cha ac e s a he gi en posi ion.
In he ideal case, we wan o ind such a egion capable o
dis inguishing each isola e. Fo ha eason, he cu en me hod
o en opy calcula ion in one posi ion (Nyk yno a e al., 2019)
in he sequence’s alignmen canno be used, as i can only
dis inguish ou a ian s because ou nucleo ides (A, C, G, T)
a e p esen ed. In his a icle, a new unique app oach using wo d
en opy is p esen ed, whe e he wo d is a nucleo ide sequence o
de ined leng h.
As e e ence-based genome assembly is no pe ec du ing
consensus calling ambiguous nucleo ides can appea in he
consensus sequences, and hey a i icially inc ease he en opy
alue as shown in Figu e 1. To p e en de ec ing alsely a iable
sequences, he whole columns o aligned sequences whe e
ambiguous nucleo ides occu we e emo ed and o p ese e he
o iginal leng h o he c ea ed consensus sequences eplaced by
gaps. Also, he de ec ed ma ke s should occu in all analyzed
isola es. Thus, when gaps appea in some sequences, he gaps
ep esen ed by dashes a e added o he whole columns as is also
shown in Figu e 1.
In he window w, whose size co esponds o wo d leng h, he
numbe o unique wo ds is de e mined, and hen he en opy o
wo ds Hwis calcula ed as
Hw= − X(kw
n·(log2
kw
n)), (3)
F on ie s in Mic obiology | www. on ie sin.o g 2Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
FIGURE 1 | Example o sequence p ep ocessing. (A) Example o aligned sequences be o e p ep ocessing and hei en opy o wo d alues o wo ds o leng h i e.
(B) Aligned sequences a e p ep ocessing and co esponding en opy o wo d alues.
FIGURE 2 | Scheme o en opy o wo ds’ calcula ion o a se o sequences a e p ep ocessing and o wo d leng h o 5 bp and one nucleo ide window shi .
whe e kwis he numbe o occu ences o unique wo ds in each
alignmen in he cu en window, and nis he numbe o isola es.
The sum is applied o e he alignmen in he cu en window.
A e es ima ing he en opy alue, he window mo es wi h a shi
o one nucleo ide. The p inciple o en opy o wo ds’ calcula ion
is shown in Figu e 2.
Fo de ined sliding window sizes w(also called wo d leng h),
he en opy signal wi h leng h N−w+1 is ob ained, whe e Nis
he analyzed genome’s leng h.
In he signal, he maximum alue is iden i ied. In heo y,
he maximum alue can be de ined as a bina y loga i hm
o he numbe o sequences. Based on he maximum alue
F on ie s in Mic obiology | www. on ie sin.o g 3Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
ha occu s in he signal, he h eshold is se as a maximum
alue educed by 0.4 bi , and signal peaks abo e he h eshold
a e labeled as po en ial gene ic ma ke s. The s a ed alue
was chosen empi ically based on he numbe o po en ial
a iable egions o ensu e ha some egions will ha e su icien
disc imina o y powe ; howe e , i can be adjus ed o analyzed
da a by he use . The aligned sequences co esponding o
peaks a e ex ac ed and analyzed. I wo peaks a e dis an
less han he window leng h, he sequences co esponding
o he peaks a e me ged and analyzed as one. De eloped
sou ce codes and unc ions in Ma lab R2020a can be ound
a h ps://gi hub.com/ma ke anyk yno a/ma ke s_de ec ion.
2.4. Phylogene ic Analysis
In he i s pa , each se o aligned sequences co esponding o
peaks was analyzed. The e olu iona y dis ances we e calcula ed
based on he Kimu a (Kimu a, 1980) model, which is used as
s anda d. F om ob ained dis ances, he phylogene ic ees we e
cons uc ed based on UPGMA. Clus e analysis was conduc ed,
and he clus e s om he ees we e es ablished and compa ed
o esul s ob ained om cgMLST. The h eshold o clus e s
de e mina ion was se o in e sec he ee so ha he maximum
numbe o clus e s wasob ained.
The esul s showed ha no se had su icien disc imina o y
powe o dis inguish be ween s ains. Tha is why he
combina ion o wo aligned sequences se s was used. The
iden i ied a iable se s o sequences we e me ged, which c ea ed
one long sequence o each analyzed bac e ial isola e. Then
o me ged sequences, again he e olu iona y dis ances we e
calcula ed and based on hem, he phylogene ic ees we e
cons uc ed and examined. The leng h o a a iable se o
sequences had a p opo ional impac on he c ea ed phylogene ic
ee. I he sequences in one a iable se had a longe leng h,
hey had a highe impac on he ee and ice e sa, he se
o sho e sequences did no a ec he inal ee so much.
Ne e heless, sequences a iabili y has he main impac on he
phylogene ic ee.
3. RESULTS AND DISCUSSION
3.1. Sequence Type De e mina ion and
cgMLST Analysis
The inco po a ed E. coli and E. aecium cgMLST, cgSNV, and
wgSNV schemes in Ridom SeqSphe e+ so wa e (Ridom, DE)
we e used o analyze he gene ic ela edness o sequenced
bac e ial genomes. Fo E. coli, his included 3,152 a ge genes
used o gene a e a cgMLST dend og am. Fo he cgSNV
minimum spanning ee analysis, 152,212 aligned nucleo ide si es
we e analyzed. In o al, 176,629 aligned nucleo ide si es we e used
o wgSNV minimum spanning ee analysis. Fo E. aecium, he
scheme included 1,423 a ge genes o cgMLST, 7,707 aligned
nucleo ide si es o cgSNV and 10,858 aligned nucleo ide si es
o wgSNV analysis. The WGS da a we e used o de e mine
sequence ypes (STs) o bo h, E. coli and E. aecium s ains. Fo E.
coli, se en housekeeping genes (adk, umC, gy B, lcd, mdh, pu A,
ecA) and o E. aecium again se en housekeeping genes (a pA,
ddl, gdh, pu K, gyd, ps S, adk) we e analyzed. Fo E. aecium
h ee STs we e iden i ied (4 ×ST 117, 1 ×ST 18, 18 ×ST 80),
and o E. coli ele en STs we e ecognized and one ST emains
unknown (1 ×ST 69, 4 ×ST 131, 1 ×ST 95, 2 ×ST 404, 2
×ST 38, 2 ×ST 1049, 4 ×ST 58, 1 ×ST 297, 1 ×ST 517, 2
×ST 101, 1 ×ST UNW). The comple e esul s a e a ached as
Supplemen a y File 1.
3.2. Po en ial Gene ic Ma ke s Analysis
Sequences wi h a high alue o wo d en opy we e iden i ied
o wo d leng hs om 50 o 400 bp wi h a s ep o 50 bp.
T adi ional labo a o y me hods can use iden i ied ma ke s o
leng h om his ange. In he i s pa , one a iable agmen
wi h high wo d en opy alues was used o dis inguish he
s ains o E. coli and E. aecium. Examining 1,358 agmen s
wi h high a iabili y o all wo d sizes showed ha he
maximum numbe o clus e s ha can be dis inguished o
E. coli was 13. Fo E. aecium, 104 po en ial gene ic ma ke s
o all wo d sizes we e analyzed, and he maximum numbe
o dis inguishable clus e s was 5. The comple e esul s a e in
Supplemen a y File 2.
The numbe o dis inguishable clus e s using one a iable
agmen is no high enough. Fo ha eason, wo a iable
agmen s wi h wo d en opy abo e h eshold we e analyzed.
Fo all combina ions o wo possible gene ic ma ke s, he
phylogene ic ees we e cons uc ed, and he numbe o clus e s
was calcula ed. In he nex s ep, he po en ial gene ic ma ke s
wi h he highes numbe o clus e s we e analyzed. Fo all wo d
sizes, 136,661 combina ions o possible gene ic ma ke s o E. coli
and 685 o E. aecium we e examined in o al.
In case o need, he combina ion o e en mo e ma ke s
can be used; howe e , he ime o analysis will g ow as
he ime complexi y is O(nm), whe e mis he numbe o
a iable agmen s.
Fo E. coli, 29 ees we e examined which di ided he isola es
o 15 clus e s. Ten ees classi y he genomes acco ding o he
11 p ede ined clus e s om cgMLST. The esul s om cgMLST
wi h 11 ma ked sequence ypes and 11 de ined clus e s and a
cladog am c ea ed based on he combina ion o wo gene ic
ma ke s wi h 15 labeled clus e s a e shown in Figu e 3.
In each case o one gene ic ma ke , he egion which
con ained he genes o hypo he ical p o ein (NP_308436.2) and
S- o mylglu a hione hyd olase (NP_308437.1) was selec ed. In
i e ou o en cases, he in e genic egion in on o hypo he ical
p o ein was also inco po a ed. As he second gene ic ma ke ,
wo egions we e mos o en ( h ee imes) chosen. The i s one
con ained ou e memb ane p o ein (NP_313225.1) and in one
case he in e genic egion in on o he gene oo. The second
one s a ed in imb ial-like adhesin p o ein (NP_310944.1) and
ended in hypo he ical p o ein (NP_310945.1). Fou o he gene ic
ma ke s we e selec ed only one ime. Two o hem consis ed
o non-coding egion and gene [PhoH amily P-loop ATPase
(NP_309293.1) in he i s case, excinuclease U ABC subuni
U C (NP_310678.2) in he second case], o ano he ma ke
pa o he gene alyl- RNA syn he ase (NP_313262.1) was
picked, and in he las case, as a gene ic ma ke he non-
coding egion was selec ed. De ailed in o ma ion is a ailable in
F on ie s in Mic obiology | www. on ie sin.o g 4Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
FIGURE 3 | Classi ica ion o 21 E. coli genomes. (A) Co e genome MLST analysis o 21 E. coli isola es wi h p ede ined ma ked clus e s (ou e ci cles) and sequence
ypes (inne ci cles). (B) Cladog am o 21 E. coli isola es based on wo iden i ied gene ic ma ke s wi h highligh clus e s ob ained om co e genome MLST analysis
(ou e ci cle) and subg oups ob ained based on analysis o iden i ied a iable sequences (inne ci cle) c ea ed by E ol iew (Sub amanian e al., 2019).
FIGURE 4 | Classi ica ion o 23 E. aecium genomes. (A) Co e genome MLST analysis o 23 E. aecium isola es wi h p ede ined ma ked clus e s (ou e ci cles) and
sequence ypes (inne ci cles). (B) Cladog am o 23 E. aecium isola es based on wo iden i ied gene ic ma ke s wi h highligh clus e s ob ained om co e genome
MLST analysis (ou e ci cle) and subg oups ob ained based on analysis o iden i ied a iable sequences (inne ci cle) c ea ed by E ol iew (Sub amanian e al., 2019).
Supplemen a y File 2 and all phylogene ic ees ha co ec ly
classi ied genomes a e depic ed in Supplemen a y Figu e 1.
Fo E. aecium, 22 ees which di ided he isola es in o nine
clus e s we e examined. I was ound ha six ees classi y he
genomes acco ding o i e p ede ined clus e s om cgMLST. In
Figu e 4, he esul s om cgMLST analysis wi h h ee labeled
sequence ypes, i e clus e s and a cladog am cons uc ed based
on he iden i ied combina ion o wo gene ic ma ke s wi h nine
labeled clus e s shown.
As in he p e ious case, in all possible combina ions,
one gene ic ma ke always emained he same, and i was
RNA me hyl ans e ase (WP_002294889.1). In one o six cases,
he in e genic egion a ound he gene was also inco po a ed
in o he a iable gene ic pa . Mg- ansloca ing P- ype ATPase
(WP_002304253.1) was in mos cases ( h ee imes) selec ed as
he second ma ke . The egion ha con ained he gene o
DUF1430 domain-con aining p o ein (WP_002288541.1) was
chosen wice. Fo he las possible gene ic ma ke , he gene o
nucleoside-diphospha e kinase (WP_002293327.1) was picked.
The comple e esul s a e a ailable in Supplemen a y File 2 and
all phylogene ic ees co ec ly classi ying genomes a e shown in
Supplemen a y Figu e 2.
3.3. Wo d Sizes
Fo de ined wo d sizes, a maximum numbe o co ec ly classi ied
clus e s o phylogene ic ees based on wo ma ke combina ions
F on ie s in Mic obiology | www. on ie sin.o g 5Feb ua y 2021 | Volume 12 | A icle 631605

Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
FIGURE 5 | Analysis o wo d size impac on numbe o co ec ly classi ied clus e s. (A) Maximum numbe o co ec ly classi ied clus e s o combina ions o wo
a iable agmen s o di e en wo d sizes and numbe o wo a iable agmen combina ions ha co ec ed classi ied analyzed E. aecium isola es. (B) Maximum
numbe o co ec ly classi ied clus e s o combina ion o wo a iable agmen s o di e en wo d sizes and numbe o wo a iable agmen combina ions ha
co ec ed classi ied analyzed E. coli isola es.
was es ablished, as is shown in Figu e 5. Fo analyzed ees he
numbe o clus e s was de e mined and he co ec genomes
classi ica ion was con olled. The esul s showed ha o E.
aecium he wo d leng hs do no signi ican ly a ec he maximum
numbe o co ec ly es ablished clus e s, which is nine o all
wo d sizes excep he leng h 50 and 300 bp. On he con a y, he
esolu ion powe o E. coli g ows wi h inc easing wo d leng hs
and he bes esul s (dis inguish 15 clus e s) can be ob ained o a
window o size 300 bp and mo e.
3.4. Resul s Ve i ica ion
The analyzed da ase s we e ex ended by comple e genomes
sequences o E. coli and E. aecium ob ained om he
GenBank da abase. Fo each bac e ium, six sequences we e
downloaded, and he STs we e de e mined. Fo E. coli
isola es (CP050212.1, CP050219.1, CP050211.1, CP050201.1,
CP050214.1, CP050205.1) i e STs we e iden i ied (1 ×ST 2705,
1×ST 2280, 1 ×ST 70, 1 ×ST 602, 2 ×ST 131), and
o E. aecium isola es (CP040236.1, CP036151.1, CP027501.1,
CP035666.1, CP035660.1, CP035220.1) one ST was ecognized,
and one ST emains unknown (5 ×ST 80, 1 ×ST UNW).
The esul s a e in Supplemen a y File 1. Using he BLAST+
( 2.6.0+, Camacho e al., 2009) he iden i ied a iable ma ke s
we e loca ed in hese genomes. The ma ke s sequences o
bo h se s (o iginal and da abase) we e analyzed oge he . The
sequences we e aligned, e olu iona y dis ances we e calcula ed,
and phylogene ic ees we e cons uc ed. To ally, 27 isola es o
E. coli belonging o nine STs we e classi ied o 17 clus e s and
29 isola es o E. aecium belonging o ou STs we e i e imes
classi ied o 11 clus e s and once o 10 clus e s. All c ea ed
phylogene ic ees a e depic ed in Supplemen a y Figu e 3.
The iden i ied gene ic ma ke s can be used o yping no el
isola es. Howe e , i a la ge numbe o new isola es should
be analyzed, he algo i hm should be un again o ind
ma ke s wi h he highes disc imina o y powe o he analyzed
bac e ial popula ion.
3.5. Analyzed Bac e ia
The p oposed app oach was es ed on wo bac e ia—E. aecium
and E. coli which belong among signi ican HAIs pa hogens.
E. aecium is a G am-posi i e bac e ium and can be a sou ce
o nosocomial in ec ions, especially in immunocomp omised
pa ien s, and causes endoca di is, bac e emia o u ina y ac
in ec ion. In addi ion, some E. aecium isola es can de elop d ug
esis ance o se e al an ibio ics g oups such as glycopep ides,
be a-lac ams, luo oquinolones, o aminoglycosides (Yoong
e al., 2004; Cas illo-Rojas e al., 2013). Ano he pa hogen
om G am-nega i e bac e ia is E. coli, which can be ound
in human in es inal lo a as a commensal. Howe e , as an
oppo unis ic pa hogen E. coli ep esen s a huge public heal h
p oblem. I is one o he mos common causes o HAIs
ypically connec ed o meningi is, pneumonia, u ina y ac
in ec ions, and so issue in ec ions (Jau eguy e al., 2008). E.
coli s ains ha e he abili y o accumula e esis ance genes; hus,
he esis ance o b oad-spec um cephalospo ins, ca bapenems,
aminoglycoside, luo oquinolones, and polymyxins is o en
obse ed (Poi el e al., 2018).
F on ie s in Mic obiology | www. on ie sin.o g 6Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
4. CONCLUSION
As yping is c ucial o iden i ying he sou ce o in ec ion
and ou b eak moni o ing, i is necessa y o ha e a highly
sensi i e yping me hod o dis inguish closely ela ed bac e ial
s ains. Howe e , cu en ly used me hods ha e insu icien
disc imina o y powe excep whole genome sequencing.
Ne e heless, WGS is ime-consuming, and da a p ocessing is
no possible in ou ine clinical p ac ice. Fo his eason, a new
app oach o de ec a iable gene ic ma ke s om WGS da a is
p esen ed. The WGS is only used once o ob ain he i s se
o da a o gene ic ma ke s de ec ion. To loca e highly a iable
sequences, wo d en opy is employed. The loca ed ma ke s can
be used o yping a closely ela ed bac e ial popula ion ia
s anda d labo a o y me hods. Thus, ano he genome sequencing
will be necessa y only i a la ge numbe o new isola es should
be analyzed.
Using he en opy-based app oach, new gene ic ma ke s wi h
he same o e en highe disc imina o y powe han cgMLST can
be iden i ied. As he leng h o wo ds om 50 o 400 bp wi h a
s ep o 50 bp is applied, po en ial ma ke s can be used in s anda d
labo a o y echniques because hey ha e sui able leng h. The only
equi emen o he p oposed app oach is ha genomes mus be
aligned, which can cause a p oblem mainly when la ge genomes
should be examined. The p esen ed me hod was es ed on bo h
G am-posi i e and G am-nega i e bac e ia.
As in bo h es ed da ase s, one ma ke wi h su icien
disc imina o y powe o yping o low di e si y bac e ial
popula ion was no ound, he combina ion o wo ma ke s
was employed. Based on wo iden i ied ma ke s, we we e able
o dis inguish he G am-nega i e and G am-posi i e bac e ial
popula ion a he same and e en a highe disc imina ion le el
han ia classic labo a o y me hods using only wo a iable
sequences ins ead o se en housekeeping genes. The iden i ied
ma ke s will be used in ou ine labo a o y me hods o bac e ial
yping. The low numbe o a iable agmen s ha should be
examined will sho en he ime o analysis.
Commonly used yping me hods mainly use coding egions
and analyze genes conse ed and p esen ed in all isola es
o he gi en species. Hence he a iabili y a e is lowe and
yping he closely ela ed popula ion is no possible. The
p oposed me hod uses he whole genome ange; hus, also
in e genic genome egions a e analyzed as hey ha e highe
a iabili y han coding a eas and can con ain gene ic ma ke s
o yping ela ed isola es. Based on iden i ied ma ke s, he
closely ela ed bac e ial popula ion can be di e si ied and o
ind ou he ansmission ou es and sou ce o ou b eaks will
be possible.
DATA AVAILABILITY STATEMENT
The da ase s p esen ed in his s udy can be ound in online
eposi o ies. The names o he eposi o y/ eposi o ies and
accession numbe (s) can be ound a : h ps://www.ncbi.nlm.nih.
go /, PRJNA675431.
AUTHOR CONTRIBUTIONS
MN, VB, KS, and HS con ibu ed o he concep ion and design
o he s udy. MN implemen ed he algo i hm and e alua ed he
esul s. ML and MB ensu ed he biological aspec s o he p ojec .
All au ho s ead and app o ed he inal manusc ip .
FUNDING
This wo k was suppo ed by a g an p ojec om he Czech
Science Founda ion (GA17-01821S).
SUPPLEMENTARY MATERIAL
The Supplemen a y Ma e ial o his a icle can be ound
online a : h ps://www. on ie sin.o g/a icles/10.3389/ micb.
2021.631605/ ull#supplemen a y-ma e ial
Supplemen a y File 1 | MLST analysis esul s o E.coli and E. aecium.
Supplemen a y File 2 | Po en ial and inal gene ic ma ke s o E. coli
and E. aecium.
Supplemen a y Figu e 1 | Phylogene ic ees ha co ec ly classi ied E. coli
genomes wi h ma ked clus e s.
Supplemen a y Figu e 2 | Phylogene ic ees ha co ec ly classi ied E. aecium
genomes wi h ma ked clus e s.
Supplemen a y Figu e 3 | Phylogene ic ees o o iginal and da abase genomes
o E. aecium and E. coli genomes wi h ma ked clus e s.
REFERENCES
A e ian, H., Vogel, M., Kwe ka , A., and Ha mann, M. (2016). Economic
e alua ion o in e en ions o p e en ion o hospi al acqui ed in ec ions: a
sys ema ic e iew. PLoS ONE 11:e0146381. doi: 10.1371/jou nal.pone.0146381
Bezdicek, M., Nyk yno a, M., Ple o a, K., B helo a, E., Kocmano a, I.,
Sedla , K., e al. (2019). Applica ion o mini-MLST and whole genome
sequencing in low di e si y hospi al ex ended-spec um be a-lac amase
p oducing Klebsiella pneumoniae popula ion. PLoS ONE 14:e0221187.
doi: 10.1371/jou nal.pone.0221187
Bolge , A. M., Lohse, M., and Usadel, B. (2014). T immoma ic: a lexible
imme o Illumina sequence da a. Bioin o ma ics 30, 2114–2120.
doi: 10.1093/bioin o ma ics/b u170
Bushnell, B. (2014). BBMap: A Fas , Accu a e, Splice-Awa e Aligne . No. LBNL-
7065E. Be keley, CA: E nes O lando Law ence Be keley Na ional Labo a o y.
Camacho, C., Coulou is, G., A agyan, V., Ma, N., Papadopoulos, J., Beale , K., e al.
(2009). BLAST+: a chi ec u e and applica ions. BMC Bioin o ma ics 10:421.
doi: 10.1186/1471-2105-10-421
Cas illo-Rojas, G., Maza i-Hi ía , M., Ponce de León, S., Amie a-Fe nández,
R. I., Agis-Juá ez, R. A., Huebne , J., e al. (2013). Compa ison
o En e ococcus aecium and En e ococcus aecalis s ains isola ed
om wa e and clinical samples: an imic obial suscep ibili y and
gene ic ela ionships. PLoS ONE 8:e59491. doi: 10.1371/jou nal.pone.
0059491
Cla k, K., Ka sch-Miz achi, I., Lipman, D. J., Os ell, J., and Saye s, E. W. (2016).
GenBank. Nucleic Acids Res. 44, D67–D72. doi: 10.1093/na /gk 1276
Foucaul , C., La Scola, B., Lind oos, H., Ande sson, S. G. E., and Raoul ,
D. (2005). Mul ispace yping echnique o sequence-based yping o
Ba onella quin ana. J. Clin. Mic obiol. 43, 41–48. doi: 10.1128/JCM.43.1.41-4
8.2005
F on ie s in Mic obiology | www. on ie sin.o g 7Feb ua y 2021 | Volume 12 | A icle 631605
Nyk yno a e al. Highly Disc imina i e Sequence Ma ke s De ec ion
Haque, M., Sa elli, M., McKimm, J., and Abu Baka , M. B. (2018). Heal h
ca e-associa ed in ec ions-an o e iew. In ec . D ug Resis . 11, 2321–2333.
doi: 10.2147/IDR.S177247
Hayashi, T. (2001). Comple e genome sequence o en e ohemo hagic Eschelichia
coli O157:H7 and genomic compa ison wi h a labo a o y s ain K-12. DNA Res.
8, 11–22. doi: 10.1093/dna es/8.1.11
Hen i, C., Leeki cha oenphon, P., Ca le on, H. A., Radomski, N., Kaas, R. S.,
Ma ie , J.-F., e al. (2017). An assessmen o di e en genomic app oaches
o in e ing phylogeny o Lis e ia monocy ogenes. F on . Mic obiol. 8:2351.
doi: 10.3389/ micb.2017.02351
Jau eguy, F., Land aud, L., Passe , V., Diancou , L., F apy, E., Guigon, G., e al.
(2008). Phylogene ic and genomic di e si y o human bac e emic Esche ichia
coli s ains. BMC Genomics 9:560. doi: 10.1186/1471-2164-9-560
Kimu a, M. (1980). A simple me hod o es ima ing e olu iona y a es o base
subs i u ions h ough compa a i e s udies o nucleo ide sequences. J. Mol. E ol.
16, 111–120. doi: 10.1007/BF01731581
Lam, M. M. C., Seemann, T., Bulach, D. M., Gladman, S. L., Chen, H., Ha ing, V.,
e al. (2012). Compa a i e analysis o he i s comple e En e ococcus aecium
genome. J. Bac e iol. 194, 2334–2341. doi: 10.1128/JB.00259-12
Li, H. (2013). Aligning sequence eads, clone sequences and assembly con igs wi h
BWA-MEM. a Xi [P ep in ]. a Xi :1303.3997
Li, H., Handsake , B., Wysoke , A., Fennell, T., Ruan, J., Home , N., e al.
(2009). The Sequence Alignmen /Map o ma and SAM ools. Bioin o ma ics 25,
2078–2079. doi: 10.1093/bioin o ma ics/b p352
Maiden, M. C. J., Byg a es, J. A., Feil, E., Mo elli, G., Russell, J. E., U win, R., e al.
(1998). Mul ilocus sequence yping: a po able app oach o he iden i ica ion o
clones wi hin popula ions o pa hogenic mic oo ganisms. P oc. Na l. Acad. Sci.
U.S.A. 95, 3140–3145. doi: 10.1073/pnas.95.6.3140
Medini, D., Dona i, C., Te elin, H., Masignani, V., and Rappuoli, R.
(2005). The mic obial pan-genome. Cu . Opin. Gene . De . 15, 589–594.
doi: 10.1016/j.gde.2005.09.006
Nyk yno a, M., Made anko a, D., Ba on, V., Bezdicek, M., Lenge o a, M., and
Sku ko a, H. (2019). “En opy-based de ec ion o gene ic ma ke s o bac e ia
geno yping,” in Lec u e No es in Compu e Science (including subse ies Lec u e
No es in A i icial In elligence and Lec u e No es in Bioin o ma ics) (Cham:
Sp inge ), 11466 LNBI, 177–188. doi: 10.1007/978-3-030-17935-9_17
Poi el, L., Madec, J.-Y., Lupo, A., Schink, A.-K., Kie e , N., No dmann, P.,
e al. (2018). “An imic obial esis ance in Esche ichia coli,” in An imic obial
Resis ance in Bac e ia om Li es ock and Companion Animals, eds S.
Schwa z, L. M. Ca aco, and J. Shen (Washing on, DC: Ame ican Socie y o
Mic obiology), 289–316. doi: 10.1128/mic obiolspec.ARBA-0026-2017
Ruppi sch, W. (2016). Molecula yping o bac e ia o epidemiological
su eillance and ou b eak in es iga ion/Molekula e Typisie ung on
Bak e ien ü die epidemiologische Übe wachung und Ausb uchsabklä ung.
J. Land Manage. Food En i on. 67, 199–224. doi: 10.1515/boku-20
16-0017
Schmi , A. O., and He zel, H. (1997). Es ima ing he en opy o DNA sequences.
J. Theo . Biol. 188, 369–377. doi: 10.1006/j bi.1997.0493
Schneide , T. D., and S ephens, R. M. (1990). Sequence logos: a new
way o display consensus sequences. Nucl. Acids Res. 18, 6097–6100.
doi: 10.1093/na /18.20.6097
Scholz, C. F. P., Jensen, A., Lomhol , H. B., B üggemann, H., and Kilian, M.
(2014). A no el high- esolu ion single locus sequence yping scheme o
mixed popula ions o P opionibac e ium acnes in i o. PLoS ONE 9:e104199.
doi: 10.1371/jou nal.pone.0104199
Schü ch, A., A edondo-Alonso, S., Willems, R., and Goe ing, R. (2018). Whole
genome sequencing op ions o bac e ial s ain yping and epidemiologic
analysis based on single nucleo ide polymo phism e sus gene-by-gene-based
app oaches. Clin. Mic obiol. In ec . 24, 350–354. doi: 10.1016/j.cmi.2017.12.016
Schwa z, D. C., and Can o , C. R. (1984). Sepa a ion o yeas ch omosome-
sized DNAs by pulsed ield g adien gel elec opho esis. Cell 37, 67–75.
doi: 10.1016/0092-8674(84)90301-5
Shannon, C. E. (1948). A ma hema ical heo y o communica ion. Bell Sys . Techn.
J. 27, 379–423. doi: 10.1002/j.1538-7305.1948. b01338.x
Sku ko a, H., Vi ek, M., Bezdicek, M., B helo a, E., and Lenge o a, M. (2019).
Ad anced DNA inge p in geno yping based on a model de eloped om eal
chip elec opho esis da a. J. Ad . Res. 18, 9–18. doi: 10.1016/j.ja e.2019.01.005
Sub amanian, B., Gao, S., Le che , M. J., Hu, S., and Chen, W.-H. (2019). E ol iew
3: a webse e o isualiza ion, anno a ion, and managemen o phylogene ic
ees. Nucl. Acids Res. 47, W270-W275. doi: 10.1093/na /gkz357
Yoong, P., Schuch, R., Nelson, D., and Fische i, V. A. (2004). Iden i ica ion o a
b oadly ac i e phage ly ic enzyme wi h le hal ac i i y agains an ibio ic- esis an
En e ococcus aecalis and En e ococcus aecium. J. Bac e iol. 186, 4808–4812.
doi: 10.1128/JB.186.14.4808-4812.2004
Con lic o In e es : The au ho s decla e ha he esea ch was conduc ed in he
absence o any comme cial o inancial ela ionships ha could be cons ued as a
po en ial con lic o in e es .
Copy igh © 2021 Nyk yno a, Ba on, Sedla , Bezdicek, Lenge o a and Sku ko a.
This is an open-access a icle dis ibu ed unde he e ms o he C ea i e Commons
A ibu ion License (CC BY). The use, dis ibu ion o ep oduc ion in o he o ums
is pe mi ed, p o ided he o iginal au ho (s) and he copy igh owne (s) a e c edi ed
and ha he o iginal publica ion in his jou nal is ci ed, in acco dance wi h accep ed
academic p ac ice. No use, dis ibu ion o ep oduc ion is pe mi ed which does no
comply wi h hese e ms.
F on ie s in Mic obiology | www. on ie sin.o g 8Feb ua y 2021 | Volume 12 | A icle 631605