RESEARCH ARTICLE
Towa ds he cha ac e iza ion o he hidden
wo ld o small p o eins in S aphylococcus
au eus, a p o eogenomics app oach
S ephan Fuchs
1
, Ma in KucklickID
2,3
, E ik LehmannID
2,3
, Alexande BeckmannID
2,3
,
Maya WilkensID
1,2,3
, Baban Kol e
4
, Ay en Mus a aye a
2,3
, Tobias Ludwig
2,3
,
Mau ice DiwoID
2,3
, Jose Wissing
5
, Lo ha Ja
¨nsch
5
, Ch is ian H. Ah ensID
6
,
Zoya Igna o aID
4
, Susanne EngelmannID
2,3
*
1Robe Koch Ins i u e, Me hodenen wicklung und Fo schungsin as uk u (MF), Be lin, Ge many,
2Uni e si y o Technical Sciences B aunschweig, Ins i u e o Mic obiology, B aunschweig, Ge many,
3Helmhol z Cen e o In ec ion Resea ch GmbH, Mic obial P o eomics, B aunschweig, Ge many,
4Uni e si y o Hambu g, Ins i u e o Biochemis y and Molecula Biology, Hambu g, Ge many, 5Helmhol z
Cen e o In ec ion Resea ch GmbH, Cellula P o eomics, B aunschweig, Ge many, 6Ag oscope, Resea ch
G oup Molecula Diagnos ics, Genomics and Bioin o ma ics & SIB Swiss Ins i u e o Bioin o ma ics, Basel,
Swi ze land
*Susanne.Engelm[email p o ec ed]
Abs ac
Small p o eins play essen ial oles in bac e ial physiology and i ulence, howe e , au o-
ma ed algo i hms o genome anno a ion a e o en no ye able o accu a ely p edic he co -
esponding genes. The accu acy and eliabili y o genome anno a ions, pa icula ly o small
open eading ames (sORFs), can be signi ican ly imp o ed by in eg a ing p o ein e idence
om expe imen al app oaches. He e we p esen a highly op imized and lexible bioin o ma -
ics wo k low o bac e ial p o eogenomics co e ing all s eps om (i) gene a ion o p o ein
da abases, (ii) da abase sea ches and (iii) pep ide- o-genome mapping o (i ) isualiza ion
o esul s. We used he wo k low o iden i y high quali y pep ide spec um ma ches (PSMs)
o small p o eins (�100 aa, SP100) in S aphylococcus au eus Newman. P o ein ex ac s
om S.au eus we e subjec ed o di e en expe imen al wo k lows o p o ein diges ion and
p e ac iona ion and measu ed wi h highly sensi i e mass spec ome e s. In o al, 175 p o-
eins wi h up o 100 aa (SP100) we e iden i ied. Ou o hese 24 ( anging om 9 o 99 aa)
we e no el and no con ained in he used genome anno a ion.144 SP100 a e highly con-
se ed and we e ound in a leas 50% o he publicly a ailable S.au eus genomes, while
127 a e addi ionally conse ed in o he s aphylococci. Almos hal o he iden i ied SP100
we e basic, sugges ing a ole in binding o mo e acidic molecules such as nucleic acids o
phospholipids.
Au ho summa y
Con en ional au oma ic genome anno a ion algo i hms o en neglec open eading
ames smalle han 300 nucleo ides (sORF). The e a e se e al easons hinde ing
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 1 / 26
a1111111111
a1111111111
a1111111111
a1111111111
a1111111111
OPEN ACCESS
Ci a ion: Fuchs S, Kucklick M, Lehmann E,
Beckmann A, Wilkens M, Kol e B, e al. (2021)
Towa ds he cha ac e iza ion o he hidden wo ld o
small p o eins in S aphylococcus au eus, a
p o eogenomics app oach. PLoS Gene 17(6):
e1009585. h ps://doi.o g/10.1371/jou nal.
pgen.1009585
Edi o : Kai Papen o , F ied ich-Schille -Uni e si a
Jena, GERMANY
Recei ed: No embe 24, 2020
Accep ed: May 7, 2021
Published: June 1, 2021
Pee Re iew His o y: PLOS ecognizes he
bene i s o anspa ency in he pee e iew
p ocess; he e o e, we enable he publica ion o
all o he con en o pee e iew and au ho
esponses alongside inal, published a icles. The
edi o ial his o y o his a icle is a ailable he e:
h ps://doi.o g/10.1371/jou nal.pgen.1009585
Copy igh : ©2021 Fuchs e al. This is an open
access a icle dis ibu ed unde he e ms o he
C ea i e Commons A ibu ion License, which
pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided he o iginal
au ho and sou ce a e c edi ed.
Da a A ailabili y S a emen : The mass
spec ome y p o eomics da a ha e been deposi ed
o he P o eomeXchange Conso ium (h p://
au oma ic anno a ion and p edic ion o sho genes: (i) sORFs possess insu icien
sequence in o ma ion o domain and homology sea ch, (ii) only a limi ed numbe o
expe imen ally alida ed sORFs can se e as empla es, and (iii) sORFs show he endency
o be species-speci ic. We hus es ablished a p o eogenomics wo k low, which is execu ed
by wo open sou ce ools, Sal and Peppe (h ps://gi lab.com/s. uchs/peppe ), and uses
pep ide da a ob ained by mass spec ome y o iden i ica ion o genes in bac e ia ha a e
ha dly p edic able by au oma ic anno a ion algo i hms. As a p oo o concep , we selec ed
S aphylococcus au eus, one o he mos equen ly sequenced bac e ia and iden i ied 36
p o eins no ye conside ed in he used genome anno a ion o S.au eus Newman. 24 he e
o a e no el small p o eins wi h up o 100 aa (SP100) in S.au eus Newman. This clea ly
demons a es ha ou wo k low is ideally sui ed o imp o e gene anno a ion o al eady
anno a ed bac e ial genomes. In he u u e, i may also acili a e p o ein and ORF de ec-
ion in no anno a ed bac e ial genomes.
In oduc ion
S aphylococcus au eus is a G am-posi i e human pa hogen o g ea clinical impo ance. S.
au eus causes mainly nosocomial in ec ions in immunocomp omized pa ien s, which a e e-
quen ly associa ed wi h di icul o ea mul id ug- esis an S.au eus pheno ypes [1]. Wi h
11,809 genome sequences (including 576 comple e genomes), which a e publicly a ailable in
he e e ence sequence da abase o he Na ional Cen e o Bio echnology In o ma ion (Re Seq;
s a us 2020-08-19), S.au eus is among he mos equen ly sequenced bac e ia. The numbe o
anno a ed open eading ames anges om 2,411 o 3,147 pe comple e genome sequence.
The en i e pan-genome o S.au eus has no ye been desc ibed, due o he ac ha he geno-
mic di e si y o S.au eus is e y high [2,3]. Howe e , a p elimina y S.au eus pan-genome
based on he compa ison o 64 S.au eus genome sequences is composed o 7,411 genes, o
which abou 20% a e conse ed cons i u ing he co e-genome [3]. The highes a iabili y has
been ound among genes coding o ex acellula and su ace-associa ed p o eins [4] which is
o pa icula impo ance as hese p o eins a e essen ially in ol ed in di ec in e ac ions wi h
he hos en i onmen du ing in ec ion.
The p o ein in en o y o se e al S.au eus s ains has been desc ibed using highly sensi-
i e mass spec ome y (MS) echniques combined wi h liquid ch oma og aphy (LC) [5–7].
Fo S.au eus s ain COL, mo e han 1,700 p o eins (abou 60% o he heo e ical p o eome)
ha e been iden i ied, quan i ied and assigned o a ious subcellula localiza ions [5,7,8],
which can help o p edic unc ions o co-exp essed and/o co-localized p o eins [9]. How-
e e , one g oup o p o eins was highly unde ep esen ed in he S.au eus p o eome: e y
small p o eins no longe han 100 amino acids (aa) (= SP100). Al oge he , Beche and col-
leagues [5] de ec ed 82 anno a ed SP100 o which only ou p o eins we e below 50 aa in
leng h (= SP50).
The expe imen al de ec ion o SP100 by sho gun p o eomics is di icul and addi ionally
hampe ed by he ac ha he co esponding sho open eading ames (sORFs) a e o en
o e looked by con en ional genome anno a ion algo i hms. The e a e se e al easons hinde -
ing au oma ed p edic ion and accu a e anno a ion o sORFs, such as insu icien sequence
in o ma ion o domain and homology sea ches, a limi ed numbe o expe imen ally alida ed
empla es, and hei endency o species speci ici y [10–12]. Hence, di e en ia ion be ween
sORFs wi h low and high coding po en ial is challenging and he numbe o alse posi i es
among p edic ed sORFs is ex emely high [13]. Gi en hese ac s, genome anno a ions
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 2 / 26
p o eomecen al.p o eomexchange.o g) ia he
PRIDE pa ne eposi o y (Vizcaino JA, Cso das A,
del-To o N, Dianes JA, G iss J, La idas I, e al.
2016 upda e o he PRIDE da abase and i s ela ed
ools. Nucleic Acids Res. 2016;44: D447-56.) wi h
he da ase iden i ie PXD017932. High Quali y MS/
MS spec a o he iden i ied yp ic pep ides unique
o SP100 a e accessible in Supplemen al
Ma e ials. Ribosome p o iling da a ha e been
deposi ed wi hin Gene Exp ession Omnibus (GEO)
unde accession numbe GSE150601.
Funding: This wo k was unded by he Deu sche
Fo schungsgemeinscha (h ps://www.d g.de/)
(GRK PROCOMPAS) o SE, by he Deu sche
Fo schungsgemeinscha (GRK PROCOMPAS) o
LJ, by he Deu sche Fo schungsgemeinscha
(INST 188/365-1 FUGG DFG) o SE, by he
Schweize ische Na ional onds zu Fo¨ de ung de
Wissenscha lichen Fo schung (h p://www.sn .ch/
de/Sei en/de aul .aspx) (197391) o CHA, by he
Deu sche Fo schungsgemeinscha (IG 73/16-1
SPP 2002) o ZI. The unde s had no ole in s udy
design, da a collec ion and analysis, decision o
publish, o p epa a ion o he manusc ip .
Compe ing in e es s: The au ho s ha e decla ed
ha no compe ing in e es s exis .
ou inely used a bi a y cu -o s o a minimum ORF leng h o 50 o 100 codons. In addi ion,
he low molecula weigh o hese p o eins complica es expe imen al isola ion and educes he
numbe o MS-compa ible pep ides.
O e he las yea s, a ious a emp s ha e been made ha add ess one o bo h o hese
issues. This includes expe imen al app oaches such as ibosome p o iling and p o eogenomics
o iden i y his g oup o p o eins as well as bioin o ma ics app oaches o a mo e eliable p e-
dic ion and comp ehensi e anno a ion o sORFs [14–26]. Fo ins ance, di e en compu a-
ional app oaches ha e been de eloped o sORF p edic ion, which ha e in common ha he
coding po en ial o a pu a i e sORF is sco ed based on one o mo e ea u es such as nucleo ide
composi ion, synonymous and non-synonymous subs i u ion a es, phylogene ic conse a ion
o p o ein domain de ec ion [13,16].
Despi e hese majo challenges, he e is no doub ha small p o eins play a pi o al ole in
essen ial cellula p ocesses, hence, i is ex emely impo an o imp o e ou abili y o unco e
his pool o hidden p o eins [20,27–30]. Fi s unc ional cha ac e iza ions p o e hei in ol e-
men in a ious cellula p ocesses such as p o ein olding, egula ion o gene exp ession, mem-
b ane anspo , p o ein modi ica ion and signal ansduc ion in di e en bac e ia ( o an
o e iew see [30]). In addi ion, some small p o eins ha e an ex acellula unc ion and exhibi
oxic o an imic obial ac i i y. In e es ingly, mos o he small p o eins cha ac e ized so a a e
associa ed wi h he cell memb ane and a e poo ly conse ed a he sequence le el [29]. In S.
au eus, he mos p ominen small p o eins a e phenol soluble modulins wi h a leng h o 20 o
40 aa [31] and del a-hemolysin (26 aa) [32]. Phenol soluble modulins possess mul iple oles in
S.au eus pa hogenesis by inducing cell lysis o blood cells, s imula ing in lamma o y esponses
and in luencing bio ilm o ma ion ( o e iew see [33]). Del a-hemolysin in e ac s wi h mem-
b anes o a ious blood cells, which concen a ion dependen ly esul s in a memb ane dis u -
bance and e en in cell lysis ( o e iew see [34]). While bo h phenol soluble modulins and
del a-hemolysin ha e been s udied in de ail in ecen yea s, da a on he iden i ica ion and
unc ional cha ac e iza ion o o he SP50 in S.au eus a e almos comple ely missing.
The a ailabili y o nume ous S.au eus genome sequences de ines i as a well-sui ed model
bac e ium o p edic ion and iden i ica ion o small p o eins and pep ides. The numbe o
anno a ed coding sequences wi h up o 303 nucleo ides is highly a iable in he 576 comple e
S.au eus genome sequences anging om 287 o 621 SP100. This is mainly a ibu ed o he
ac ha a ious algo i hms o genome anno a ion we e applied. Among S.au eus e e ence
s ains, s ain Newman plays a pi o al ole. Fi s isola ed in 1952 om a human in ec ion [35],
i is one o he mos equen ly used S.au eus s ains in in ec ion models as i is cha ac e ized
by a ela i ely s able pheno ype. In addi ion, ou p ophages we e iden i ied in he genome o
s ain Newman, inse ed a di e en si es in he ch omosome, exceeding he egula ly
obse ed numbe o p ophages in S.au eus [36–38]. In a mu ine in ec ion model, he loss o
all ou p ophages signi ican ly educed he i ulence po en ial o he s ain [39]. The exis ence
o hese p ophages made S.au eus Newman an excellen model o s udy hei impac on i u-
lence and cells physiology. The genome sequence o s ain Newman, was p edic ed o encode
a leas 2,854 p o eins, again, he numbe o anno a ed SP100 is a he low and amoun s o 493
p o eins [40] (NC_009641.1; genome anno a ion om 2020-02-17).
To mo e comp ehensi ely iden i y p o eins wi h up o 100 aa, we used S.au eus Newman
as a model sys em and de eloped an ully- ea u ed p o eogenomics wo k low ha combines in
silico ansla ion o he en i e genome sequence, a ious LC-MS/MS wo k lows, and a bioin-
o ma ics pipeline o pep idomics da a analyses. The wo k low is highly op imized, lexible
and eady o use in o he bac e ial species.
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 3 / 26
Me hods
Bac e ial s ains, cul i a ion condi ions and cell lysis
S.au eus Newman [35] was cul i a ed in 100 mL complex medium (TSB) a 37˚C and 120
pm o an op ical densi y a 540 nm (OD
540
) o 1 and 7. Cells we e ha es ed by mul iple cen-
i uga ion s eps and dis up ed by cell homogeniza ion (Fas P ep-24, MP Biomedicals) ( o
de ails see [41]). All expe imen s ha e been pe o med wi h h ee biological eplica es. The
p o ein concen a ion was de e mined using he Ro i-Nanoquan assay (Ro h, Ka ls uhe, Ge -
many) and he p o ein solu ion was s o ed a -20˚C.
F ac iona ion o p o eins and pep ides and p o eoly ic clea age
Gel-based app oach. 40 μg o cy oplasmic p o eins we e sepa a ed by one dimensional
SDS polyac ylamide gel elec opho ese (1D SDS PAGE) acco ding o Laemmli [42] wi h he
ollowing modi ica ions: he loading bu e consis ed o 3.75% ( / ) glyce ol, 1.25% ( / ) ß-
me cap oe hanol, 0.6% (w/ ) SDS, 0.0014% (w/ ) b omophenol blue, 16.5 mM T is-HCl (pH
6.8). The sepa a ion gel con ained 12% (w/ ) ac ylamide gel (wi h 0.32% bisac ylamide), 0.375
M T is-HCL (pH 8.8), 0.255% (w/ ) SDS, 0.062% (w/ ) APS, and 0.062% ( / ) TEMED and
he s acking gel 5% (w/ ) ac ylamide (wi h 0.13% (w/ ) bisac ylamide), 0.125 M T is-HCl (pH
6.8) 0.25% (w/ ) SDS, 0.075% (w/ ) APS, and 0.075% ( / ) TEMED.
P o eins we e ixed wi h 40% ( / ) e hanol and 10% ( / ) ace ic acid o one hou and sub-
sequen ly s ained wi h colloidal coomassie [43] o one hou . In-gel diges ion using ypsin
and ex ac ion o he pep ides we e ca ied ou as desc ibed by Le ch e al. [43] wi h an addi-
ional ex ac ion s ep using ace oni ile.
Fo diges ion wi h Lys-C, a bu e con aining 25 mM TRIS/HCl and 1 mM EDTA (pH 8.5)
was used. The applied enzyme concen a ion was 1/40 o he o al p o ein concen a ion.
Diges ion o AspN was pe o med in 10 mM T is-HCl (pH 8.0) wi h a inal AspN concen a-
ion o 1/50 o he o al p o ein concen a ion.
Gel- ee app oach. The gel- ee app oach was pe o med by applying yp ic in-solu ion
diges ion ollowed by an Oasis HLB-SPE-ca idge pu i ica ion and SCX- ac iona ion. In
de ail, 40 μg o c ude p o ein ex ac we e sol ed in 8 M u ea and 2 M hiou ea and adjus ed o
a inal concen a ion o 6 M u ea. A e addi ion o 1.6 μL o 5 mM DTT in 50 mM ammonium
bica bona e solu ion (pH 7.8), he p o ein solu ion was incuba ed o 30 min a oom empe a-
u e. Fo alkyla ion, 1 μL o a eshly p epa ed 55 mM IAA in 50 mM ammonium bica bona e
bu e was added o 10 μL o he p o ein solu ion and incuba ed o 20 min in he da k a
oom empe a u e. Subsequen ly, he solu ion was adjus ed o a inal concen a ion o 1 mM
CaCl
2
and 1 M u ea using CaCl
2
sol ed in a 50 mM ammonium bica bona e bu e . Fo diges-
ion, 1 μg ypsin (in 50 mM ammonium bica bona e and 1 mM CaCl
2
) was applied o 50 μg
p o ein. Diges ion was pe o med o 12 h a 37˚C wi h gen le agi a ion (50 pm) and s opped
by acidi ica ion o a pH alue o �2.5 wi h 10% o mic acid.
Fo pep ide pu i ica ion, Oasis HLB-SPE-ca idges (1cc, 10mg, Wa e s, Mil o d, MA,
USA) we e ini ially condi ioned wi h ace oni ile and hen wi h 0.5% o mic acid in 60% ace o-
ni ile. Subsequen ly, ca idges we e equilib a ed wi h wo olumes o 0.5% o mic acid. Sam-
ples we e loaded on he ca idges; he low- h ough was collec ed and again loaded on he
ca idge. The pep ides we e washed i e imes wi h 0.5% o mic acid (FA) and elu ed wice
wi h 0.85 mL 60% ACN 0.5% FA. Elua es we e d ied in a speed ac (Eppendo Concen a o
plus, Eppendo AG, Hambu g, Ge many) and ozen a -20˚C.
SCX ac iona ion was done as desc ibed by Kumme e al. [44]. To educe he numbe o
ac ions o eigh , pep ide-con aining ac ions we e combined.
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 4 / 26
Pep ide desal ing. ZipTips (C18, Me ck Millipo e, Bille ica, MA, USA) we e condi ioned
wi h 50% ace oni ile wice. Subsequen ly hey we e equilib a ed h ee imes wi h 0.1% FA in
5% ace oni ile. 10 μL o each pep ide ac ion ( esol ed in 20 μL 0.1% FA in 5% ace oni ile
o 60 min) we e loaded on he C18 ma ix o he ip by aspi a ing 10 imes. Elu ion was pe -
o med h ee imes by aspi a ing i e imes wi h 0.1% FA in 60% ace oni ile in a new mic o
es ube. Samples we e d ied in a speed ac.
Liquid ch oma og aphy coupled mass spec ome y (LC-MS/MS)
Fo LC-MS/MS, each pep ide ac ion o a sample was sol ed in 16 μL o 0.1% FA in 3% ace o-
ni ile o one hou , ul asonica ed in a wa e ba h o 5 min and ul acen i uged.
O bi ap Velos P o MS. LC-MS/MS uns wi h he O bi ap Velos P o MS (The mo
Fishe Scien i ic Inc, Wal ham, MA USA) we e done as desc ibed by Le ch and cowo ke s
[43].
O bi ap Fusion MS. LC-MS sys em and used columns a e desc ibed by Buli a and
cowo ke s [45]. A 200 min g adien was applied, s a ing wi h 3.7% bu e B (80% ace oni ile,
5% DMSO and 0.1% o mic acid) and 96.3% bu e A (0.1% o mic acid, 5% DMSO): 0–5 min
3.7% B; 5–125 min 3.7–31.3% B; 125–165 min 31.3–62.5% B; 165–172 min 62.5–90.0% B; 172–
177 min 90% B; 177–182 min 90–3.7% B, 182–200 min 3.7% B.
P ima y Scans we e pe o med a he O bi ap in he p o ile modus scanning an m/z o
350–1800 wi h a esolu ion ( ull wid h a hal maximum a m/z 400) o 120,000 and a lock
mass o 445.1200. Using he Xcalibu so wa e (The mo Fishe Scien i ic Inc., San Jose, CA,
USA), he mass spec ome e was con olled and ope a ed in he “ op speed” mode, allowing
he au oma ic selec ion o as much as possible wice o ou old-cha ged pep ides in a h ee-
second ime window, and he subsequen agmen a ion o hese pep ides. In he non- a ge ed
modus, p ima y ions (±10 ppm) we e selec ed by he quad upole (isola ion window: 1.6 m/z),
agmen ed in he ion ap using a da a dependen CID mode ( op speed mode, 3 seconds) o
he mos abundan p ecu so ions wi h an exclusion ime o 13 s and analysed by he ion ap.
P o ein da abase gene a ion
To conside i s ull coding po en ial, he genome sequence o he S.au eus subsp. au eus s ain
Newman (NC_009641.1) was ansla ed in all six eading ames om s op o s op codon
using he SALT ool (h ps://gi lab.com/s. uchs/peppe ). Genome ci cula i y has been consid-
e ed and en ies wi h less han 9 amino acids excluded esul ing in a o al o 177,532 sequence
en ies.
MS Da a analysis and s a is ics
Analyses o he ob ained MS and MS/MS da a we e pe o med using MaxQuan (Max Planck
Ins i u e o Biochemis y, Ma ins ied, Ge many, www.maxquan .o g, e sion 1.5.2.8) and he
ollowing pa ame e s: pep ide ole ance 5 ppm; a ole ance o agmen ions o 0.6 Da; a i-
able modi ica ions: me hionine oxida ion and ace yla ion a p o ein N- e minus, ixed modi i-
ca ion: ca bamidome hyla ion (Cys); a maximum o wo missed clea ages and ou
modi ica ions pe pep ide was allowed. Fo he iden i ica ion o SP100, a minimum o one
unique pep ide pe p o ein and a ixed alse disco e y a e (FDR) o 0.0001 o PSMs and 0.01
o p o eins was applied. The minimum sco e was se o 40 o unmodi ied and modi ied pep-
ides, he minimum del a sco e was se o 6 o unmodi ied pep ides and o 17 o modi ied
pep ides. All samples we e sea ched agains he S.au eus T ansla ion Da abase (TRDB) wi h a
decoy mode o e e ed sequences and common con aminan s supplied by MaxQuan
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 5 / 26
Fo iden i ica ion o non-anno a ed open eading ames based on iden i ied pep ides a
p o eogenomics ool has been de eloped ha desc ibes any iden i ied pep ide in he degene -
a ed DNA code sequence. By his, exac ma ches wi hin he e e ence genome can be ound
and, subsequen ly, il e ed based on exis ing anno a ion and loca ion.
Phylogene ic and unc ional analyses
Fo phylogene ic analyses we downloaded all comple e genome sequences o S.au eus
(n = 541) and s aphylococci (n = 165) om NCBI Re Seq (s a e o 2020-05-18). All SP100
sequences we e sea ched agains he downloaded genome sequences using blas n. Based on
he bes hi alignmen o e e y bac e ial ch omosome, he iden i y ela ed o he ull que y
leng h was calcula ed. Only alignmen s sha ing a leas 90% iden i y wi h he ull que y
sequence we e conside ed. Based on his, ela i e species and genus conse a ion a es ha e
been calcula ed.
Fo unc ional analyses we sea ched all SP100 sequences agains he eggNOG da abase 5.0
using eggNOG-mappe 2.0 (de aul pa ame e s). The axonomic scope was au oma ically
adjus ed o each que y o ensu e co ec classi ica ion o phage p o eins. Only unc ions om
one- o-one o hology we e ans e ed.
Ribosome p o iling
Lib a y p epa a ion. S.au eus cells om 30 mL cul u e g own in TSB medium o OD
550
= 1 we e ha es ed by apid cen i uga ion and esuspended in 390 μL ice cold 20 mM T is
lysis bu e pH 8.8, con aining 10 mM MgCl
2
x 6 H
2
O, 100 mM NH
4
Cl, 20 mM T is (pH 8.0),
0.4% T i on-X-100, 4 U DNase, 0.4 μL Supe ase-In (Ambion), 1mM chlo amphenicol. Cells
we e dis up ed by cell homogeniza ion (Fas P ep-24, MP Biomedicals) wi h 0.5 mL glass
beads (diame e 0.1 mm) o 30 s a 6.5 m/s ollowed by incuba ion on ice o 5 min. These
s eps we e epea ed wice. To emo e cell deb is, cell lysa es we e cen i uged and subsequen ly
s o ed a -80˚C and 100 A
260
uni s o ibosome-bound mRNA ac ion we e subjec ed o
nucleoly ic diges ion wi h 10 uni s/μl mic ococcal nuclease (The mo ishe ) in bu e wi h
pH 9.2 (10 mM T is pH 11 con aining 50 mM NH
4
Cl, 10 mM MgCl
2
, 0.2% i on X-100,
100 μg/mL chlo amphenicol and 20 mM CaCl
2
). The RNA agmen s we e deple ed using
S.au eus iboPOOL RNA oligo se (siTOOLs, Ge many) and he lib a y p epa a ion was pe -
o med as p e iously desc ibed [46].
Bioin o ma ic analyses o ibosome p o iling RNAs. Raw sequencing eads we e
immed using FASTX Toolki (quali y h eshold: 20) and adap e s we e cu using cu adap
(minimal o e lap o 1 n ) and mapped o he genome e sion NC_009641.1 (NCBI, Janua y
2020). Following ex ac ion o eads mapping o RNAs, he emaining eads we e uniquely
mapped o he e e ence genome using Bow ie, pa ame e se ings: -l 16 -n 1 -e 50 -m 1—
s a a–bes y. Non-uniquely mapped eads we e non-conside ed ( o mo e de ails see [47]).
Resul s
Es ima ing he numbe o spu ious sORFs in S.au eus Newman using
single-nucleo ide pe mu a ion es ing
A single nucleo ide pe mu a ion es was used o e i y he global con idence and signi icance
o ORFs based on hei leng h only. Fo his pu pose, ORFs we e de ec ed in he genome
sequence o S.au eus Newman (NC_009641.1; NCBI ansla ion able 11; longes ORFs p e-
e ed). As alse-posi i e es ima e, we used he median numbe o ORFs de ec ed in pe mu ed
genome sequences (n = 1000), ha show he same nucleo ide composi ion bu in andom
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 6 / 26
o de and can he e o e be assumed o no con ain any biological in o ma ion (Fig 1A).
Acco dingly, he alse disco e y a e (FDR) o ORFs de ec ed in he biological sequence is
113% o hose wi h a maximum leng h o 63 bp (coding o 20 aa), 133% o hose wi h a
leng h be ween 66 and 153 bp (21 o 50 aa) and 94% o hose wi h a leng h be ween156 and
303 bp (51 o 100 aa). This highligh s he need o addi ional e idence o he eliable anno a-
ion o sORFs. In con as , coding sequences (CDS) wi h a leng h o a leas 396 bp (�131 aa)
did no occu by chance in he pe mu ed sequence se . Since s a and s op codons end o be
AT- ich, he FDR o sORFs inc eases wi h an inc easing GC con en o an o ganism (Fig 1B).
C ea ing mo e comp ehensi e p o ein da abases o S.au eus Newman
using in silico ansla ion
To iden i y small p o eins no co e ed in he Re Seq anno a ion, we gene a ed a p o ein da a-
base conside ing he ull coding po en ial o S.au eus Newman by ansla ing all six eading
ames o he espec i e genome sequence and c ea ing a sepa a e p o ein en y o each
sequence be ween wo s op codons. The esul ing TRansla ion Da aBase (TRDB) comp ises
177,532 sequences wi h a minimum leng h o 9 aa. We es ablished an au oma ed wo k low o
he ansla ion o (ci cula ) bac e ial genome sequences (Fig 2A). The co esponding py hon-
based ool called Sal is publicly a ailable (h ps://gi lab.com/s. uchs/peppe ). I allows he
ex ac ion o all po en ial ORFs om a ci cula o linea genome sequence and suppo s bo h
s a o s op and s op o s op codon ex ac ion (acco ding o NCBI ansla ion able 11). The
ex ac ed sequences can be au oma ically ansla ed and, i equi ed, diges ed in silico in o
indi idual pep ides using p ede ined diges ion pa e ns o di e en enzymes. All DNA, p o ein
and pep ide sequences can be s o ed in indi idual FASTA iles. The espec i e sequence heade
can be ully cus omized o mee speci ic equi emen s. Mo eo e , o each sequence collec ion,
ea u e ables can be expo ed as abula o delimi ed ex iles lis ing di e en physicochemical
Fig 1. sORF equencies in biological and pe mu ed genome sequences. (A) Es ima ion o he p opo ion o alse-posi i e sORFs in p edic ions based solely on
s a and s op codons: All po en ial ORF sequences (NCBI ansla ion able 11; longes ORF a ian s p e e ed) we e ex ac ed om he genome sequence o S.
au eus Newman and 1,000 pe mu ed sequence de i a i es showing he same nucleo ide composi ion bu in andom o de and can he e o e be assumed o no longe
con ain any biological in o ma ion. Resul ing ORFs we e binned based on hei leng h. Bin sizes a e shown o he genuine e e ence sequence (o ange) and he
pe mu ed sequences (g ey; as median) up o a maximum ORF leng h o 303 bp (= 100 aa). Especially small ORFs end o occu andomly. (B) Impac o GC con en
on he numbe o spu ious ORF: Acco ding o (A), bin sizes a e gi en o genome sequences and hei pe mu ed sequence de i a i es (n = 1,000; as median) wi h
a ying GC con en . The used e e ence genome sequences we e NC_000913.3 (Esche ichia coli K-12 subs . MG1655; 50.8%GC; 4,641,652 bp), NC_000964.3
(Bacillus sub ilis subsp.sub ilis s . 168; 43.5%GC; 4,215,606 bp), NC_007633.1 (Mycoplasma cap icolum subsp.cap icolum ATCC 27343; 23.8%GC, 1,010,023 bp),
NC_009641.1 (S aphylococcus au eus subsp.au eus s . Newman; 32.9%GC; 2,878,897 bp), NC_010162.1 (So angium cellulosum So ce56; 71.4%GC; 1,303,779 bp).
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g001
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 7 / 26
Fig 2. Bac e ial p o eogenomics wo k low p o ided by Sal and Peppe . (A) C ea ion o p o ein and pep ide
da abases using Sal : Based on a FASTA ile as inpu , Sal ex ac s all po en ial ORFs om a gi en (ci cula ) genome
sequence using di e en me hods (s op o s op codon, s a o s op codon). The esul ing ORF sequences a e hen
ansla ed in silico in o p o ein sequences ha can be u he diges ed using di e en in silico p o eases (T yspin,
Chymo ypsin, Asp-N, Lys-C and P o einase K). Fo each le el (ORFs, p o eins, pep ides) indi idual FASTA iles and
ables ( ab-sepa a ed alues, TSV) a e c ea ed, lis ing a ious sequence-de i ed p ope ies such as molecula weigh o
isoelec ic poin s. (B) P o eogenomics analyses using Peppe : Pep ide o spec um ma ches (PSMs) ob ained om
di e en samples a e di ec ly ex ac ed om MaxQuan e idence iles (MQ). Spec al da a can be ex ensi ely e alua ed
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 8 / 26
p ope ies such as molecula weigh s, isoelec ic poin s o g and a e age o hyd opa hy
(GRAVY) alues. To he bes o ou knowledge, his is he only eely a ailable ool ha o e s
such a a ie y o unc ions ( u he in o ma ion see h ps://gi lab.com/s. uchs/peppe ).
C ea ing a ully au oma ed ye lexible wo k low o bac e ial
p o eogenomics
To deduce pu a i e ORFs om a lis o iden i ied pep ides, we c ea ed a ule-based expe sys-
em called Peppe which is ully au oma ed and op imized o bac e ial p o eogenomics (Fig
2B). In b ie , sequences and quali y measu es o pep ide spec um ma ches (PSMs) a e
ex ac ed om e idence iles (and, op ionally, ms/ms iles) p o ided by he MaxQuan so -
wa e (Max Planck Ins i u e o Biochemis y, Ma ins ied, Ge many, e sion 1.5.2.8; h p://
www.maxquan .o g). Di e en spec um- and quali y-based il e c i e ia can be hen applied
au oma ically o es ic he analyses o high-quali y PSMs only (a leas 5 consecu i e y- o b-
ions o a leas 2 �4 y-ions o b-ions o a leas 4 b- and 4 y-ions) (see also Table 1). The espec-
i e il e c i e ia we e deduced om a e y ex ensi e isual inspec ion and assessmen o he
MS/MS spec a by expe s. The objec i e was o educe he numbe o alse-posi i es when
applying a cu -o o only one unique pep ide pe p o ein o iden i ica ion o SP100. We
equi ed ha he e be a sequence ag o a leas i e consecu i e b o y agmen ions o wo
imes ou consecu i e b o y ions in a spec um o be conside ed [21]. In addi ion, pep ide
speci ic mass acks should ha e signi ican le els abo e he backg ound le els. Hence, Peppe
is able o au oma ically sco e and il e high-quali y PSMs on hese speci ic equi emen s and
o apply addi ional il e s ela ed o he in ensi y co e age (>0.1), And omeda Sco e (�40)
and pos e io e o p obabili y (<0.1) (Table 1). High-quali y MS/MS spec a o yp ic pep-
ides unique o SP100 in S.au eus a e accessible in he Supplemen al Ma e ials.
In he nex s ep, all po en ial coding si es (DNA ma ches) a e iden i ied o each high-qual-
i y PSM wi hin he gi en genome sequence conside ing he degene a ed na u e o he DNA
code. On he basis o he DNA ma ches ound, po en ial ORFs a e deduced acco ding o he
ollowing ules: i s , an ORF mus con ain all successi e DNA ma ches ha a e encoded on
he same s and and in he same eading ame and no sepa a ed by an in e posed s op
codon. Secondly, he ORF is ex ended un il he i s s op codon downs eam o he las DNA
ma ch and encoded in he same eading ame. In he inal s ep, p edic ing he ansla ional
s a si e, he mos ups eam DNA ma ch co e ed by he po en ial ORF plays an essen ial ole.
Th ee di e en cases can be dis inguished he e: (1) The iden i ied pep ide encoded by he
mos ups eam DNA ma ch is no a p o eoly ic (e.g. yp ic) p oduc . In his case, he i s
codon o he mos ups eam DNA ma ch is assumed o be he ansla ion s a si e. Since N-
e minal me hionine esidues a e clea ed om a numbe o bac e ial p o eins du ing
o apply di e en spec um quali y and eplica ion c i e ia can be de ined o es ic he analysis o highly- eliable PSMs
only. Respec i e coding si es a e de e mined in a gi en (ci cula ) genome sequence p o ided as FASTA ile. The
esul ing coding si es (DNA ma ches), ha can be il e ed e.g. by exclusi i y a e used o p edic he pu a i e open
eading ames. Addi ional in o ma ion such as po en ial ibosomal binding si es, gene syn eny based on he e e ence
genome anno a ion, and conse a ion in gi en sequence collec ions (p o ided as FASTA iles) a e collec ed. Di e en
iles a e c ea ed o a chi e all analysis pa ame e s (log ile), esul s on pep ide, DNA ma ch, and ORF le el (TSV), and
an upda ed e e ence genome anno a ion in eg a ing he iden i ied DNA ma ches and ORFs (Genbank ile; GB). (C)
Resul s isualiza ion: GB iles c ea ed by Peppe can be used o esul s isualiza ion using hi d-pa y so wa e (he e:
Geneious P ime, Bioma e s L d.). The genome sequence (black line) wi h coo dina es is shown on op. Exis ing
anno a ions a e highligh ed in yellow and g een. ORFs wi h he highes coding po en ial ega ding Peppe and po en ial
ibosomal binding si es (RBS) a e highligh ed in ed. DNA ma ches a e show in ligh - ed, i he espec i e pep ide is no
encoded elsewhe e in he genome (exclusi e ma ch), o ligh -blue, i mul iple coding si es exis o he espec i e
pep ide.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g002
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 9 / 26
me hionine. Fo he emaining 14 sORFs, any possible in- ame s a codon was assigned and
he esul ing ORFs we e e alua ed based on di e en c i e ia implemen by Peppe (S1 Table).
While he majo i y o he 24 new sORFs has been sugges ed o s a wi h ATG (n = 11), six
sORFs p esumably ini ia e a non-canonical ansla ional s a si es such as TTG (n = 2), ATT
(n = 1), GTG (n = 1), ATA (n = 1) and ATC (n = 1). Fo se en sORFs, ini ia ion a comple ely
unexpec ed codons was pos ula ed. In hese cases, he iden i ied pep ides encoded by he mos
ups eam DNA ma ch o he espec i e ORF was no a p o eoly ic p oduc and he i s codon
o he mos ups eam DNA ma ch was assumed o be he ansla ion s a si e (Fig 9A). To
ind u he suppo o he p edic ed sORFs, we looked o possible ups eam ibosomal bind-
ing si es. Hence, eigh o he newly de i ed sORFs a e p eceded by a ibosomal binding si e
wi hin a dis ance up o 14 nucleo ides ups eam o he pu a i e s a codon.
Based on hei genome localiza ion, he 24 no el sORFs can be classi ied in o h ee main
ca ego ies: (i) loca ed in in e genic egions, (ii) o e lapping wi h o he ORFs ei he wi h he 5´
o wi h he 3´-end bu in a di e en eading ame, and (iii) loca ed wi hin anno a ed p o ein
coding sequence bu in a di e en eading ame. Fo he la e wo we can addi ionally dis in-
guish be ween hose loca ed a he same s and and hose loca ed a he complemen a y s and.
The majo i y (n = 13) belongs o he g oup (iii) o which i e a e localized a he same s and
(Fig 9B). This g oup o p o eins is highly in e es ing and i s e idence o hei exis ence was
ecen ly epo ed o se e al o he o ganisms using N- e minomics o ibosomal p o iling in
combina ion wi h e apamulin [24,28,55,56]. Se en o he newly iden i ied sORFs we e
de ec ed by a leas wo di e en app oaches. Fou sORFs we e alloca ed o g oup (ii) and en
o g oup (i). No ably, h ee SP100 belonging o g oup (i) a e encoded by egions wi hin pseu-
dogenes (Fig 10A–10C). These a e sORFSaNew0004 (pseudogene NWMN_RS15675),
sORFSaNew0010 (pseudogene NWMN_RS01305), and sORFSaNew0044 (pseudogene
NWMN_RS04585). While o NWMN_RS01305 only one ame shi mu a ion leads o an
Fig 7. Phylogene ic conse a ion o he iden i ied SP100 a species and genus le el. SP100 we e sea ched agains he
Re Seq genome sequences using blas n. Based on he bes hi alignmen o e e y genome, he iden i y ela ed o he
ull que y leng h was calcula ed. Only alignmen s sha ing 90% iden i y wi h he ull-leng h que y sequence we e
conside ed. On he basis o hese esul s, ela i e species and genus conse a ion a es ha e been calcula ed.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g007
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 16 / 26
Fig 8. Func ional classi ica ion o he iden i ied SP100. Conse ed p o ein domains we e iden i ied using he
eggNOG da abase. On he basis o his, 140 SP100 (80%) we e success ully assigned o an eggNOG o hologues clus e
p o iding addi ional suppo o hei exis ence by biological signi icance.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g008
Fig 9. Cha ac e is ics o no ye anno a ed SP100. In o al 24 SP100 we e iden i ied in S.au eus Newman, which
we e no co e ed by he used gene anno a ion (NC 009641.1: genome anno a ion om 2020-02-17). The encoding
sORFs we e de i ed by Peppe on he basis o he iden i ied pep ides and speci ic c i e ia conce ning he ansla ional
s a codon, he Shine Dalga no sequence and he leng h o he space be ween bo h. (A) Dis ibu ion o di e en
ansla ional s a codons be ween he newly p edic ed sORFs. (B) Cha ac e is ics o he genome localiza ion o he
newly p edic ed sORFs: (i) in e genic egions, (ii) pa ly o e lapping wi h ano he ORF a he same s and o a he
complemen a y s and, and (iii) comple ely o e lapping wi h ano he ORF a he same s and o a he complemen a y
s and.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g009
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 17 / 26
in e up ion o he open eading ame and ansla ion o he ull leng h p o ein canno be
excluded, o NWMN_RS15675 and NWMN_RS01305 se e al in e up ions ha e been
de ec ed and ansla ion o he ull leng h p o ein is ex emely unlikely.
T ansla ion o iden i ied sORFs
We addi ionally pe o med ibosome p o iling o S.au eus g own unde he exac same
g ow h condi ions as used o assessing he p o eome. Ribosome p o iling p o ides a snapsho
o ansla ion [57] and he posi ion o he ansla ing ibosomes can be assessed wi h codon
Fig 10. Iden i ied sORFs localized wi hin pseudogenes. Schema ic p esen a ion o he NWMN_RS15675 ( npA) (A)
NWMN_RS01305 (B) and NWMN_RS04585 locus (C) based on he anno a ion o he S.au eus Newman genome
sequence (NC_009641.1; genome anno a ion om 2020-02-17). Anno a ed pseudogenes a e shown in ligh g een and
he de i ed coding sequence (CDS) in yellow. Ma ched unique pep ides iden i ied by MS/MS a e depic ed in da k
g een and he bes ORF de i ed by Peppe on he basis o he iden i ied unique pep ides and addi ional ea u es is
depic ed in da k ed. Peppe analyses o diges ion wi h ypsin a e p esen ed.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g010
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 18 / 26
p ecision [58]. The ob ained sequencing eads, which ep esen ibosome-p o ec ed mRNA
agmen s (RPF), we e mapped o he genome o S.au eus Newman and ansla ion p o iles
we e gene a ed ( o mo e de ails see [47]). Fo 135 (76%) sORFs iden i ied by using ou p o-
eogenomics pipeline we de ec ed RFPs. Among hem a e i e ha we e missing in he used
genome anno a ion. Fo an addi ional se en o he newly anno a ed sORFs, based on ou
app oaches, eads ha e been mapped o he espec i e coding egion close o he signal o
noise h eshold; hey would equi e o he alida ion expe imen s.
In case o sORFs embedded wi hin p o ein coding egions a he same s and, we we e no
able o clea ly assign RPFs. In Esche ichia coli, Re apamulin-enhanced Ribo-seq analysis
(Ribo-RET), which combines e apamulin speci ic a es s o ini ia ing ibosomes and ibo-
some p o iling o map ansla ional s a si es, e ealed pu a i e in e nal s a si es in a numbe
o genes [24,55]. Howe e , p elimina y a emp s o use his echnique in S.au eus o mapping
ansla ional ini ia ion si es ha e ailed so a , possibly due o an ine icien anspo o e apa-
mulin in o he cell.
27 sORFs, which we e iden i ied by he p o eogenomics app oach, showed no ansla ional
ac i i y a all. Among hem a e 15 sORFs which a e localized on h ee di e en lysogenic
phages. I is in e es ing o no e, ha hese phages a e highly simila and since we only used
uniquely mapped eads (i.e. mapping o only one posi ion o he genome), i is likely ha he
RPFs we e disca ded due o ambigui y. In sum, we ha e obse ed ansla ional ac i i y o he
majo i y (76%) o he sORFs de ec ed by ou p o eogenomics app oach, alida ing he p edic-
i e powe o he used MS da a se s and TRDB in combina ion wi h he he e de eloped p o eo-
genomics ool Peppe o da a analyses. (S4 Table).
Discussion
P edic ion and iden i ica ion o sORFs coding o p o eins smalle han 100 aa (SP100) is s ill
e y challenging om bo h compu a ional and expe imen al poin s o iew. In ecen yea s,
se e al s udies s a ed o add ess his issue in a sys ema ic way by combining compu a ional
p edic ion and expe imen al alida ion using ibosome p o iling and/o mass spec ome y
[16,17,19,22,24,48,49,59–63] P o eomics based iden i ica ion o sORFs elying on mass spec-
ome y is especially ad an ageous as i p o ides di ec e idence o he exis ence o sORF
encoded p o eins and alida es no only ansla ion o sORFs bu also s abili y o hei gene
p oduc s [64,65]. Howe e , mass spec ome y based iden i ica ion o sORFs is aced by se -
e al challenges: he e y low numbe o pep ides a ailable o mass spec ome y and he
dependence on amino acid e e ence sequences. To add ess his, we he e de eloped a highly
op imized and lexible da a analysis pipeline o bac e ial p o eogenomics, co e ing all s eps
om (i) p o ein da abase gene a ion, (ii) da abase sea ch, (iii) pep ide- o-genome mapping,
and (i ) esul isualiza ion. The wo k low is based on ou bac e ial p o eogenomics pipeline
Peppe , ex ended by Sal , a genome ansla o gene a ing p o ein and pep ide da abases using
di e en me hods (e.g., s op- o-s op o s a - o-s op ansla ion). Peppe ep esen s a ule
based expe sys em ha enables ully au oma ed MS da a analysis o empi ical and e idence
based gene anno a ion. Au oma ic ule based selec ion o high quali y PSMs o pep ide iden i-
ica ion combined wi h an imp o ed s a codon de ec ion based on N- e minal pep ides dis-
inguishes his wo k low om al eady exis ing pipelines o bac e ial p o eogenomics
[16,17,66–68]. I is hus pa icula ly well placed o iden i y exp essed SP100 in bac e ia by one
unique pep ide comple ely independen o exis ing genome anno a ions. Compa ed o e y
sensi i e ibosome p o iling app oaches, which a e mo e equen ly used o sORF iden i ica-
ion, highly con iden MS based iden i ica ion o p o eins as acili a ed by his wo k low
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 19 / 26
p o ides di ec e idence o he exis ence o small p o eins and alida es no only ansla ion
o sORFs bu also s abili y o hei gene p oduc s unde he es ed condi ions.
To e alua e commonly used expe imen al app oaches o iden i ica ion o SP100 in S.
au eus, we applied a gel-based and a gel- ee LC-MS/MS app oach in combina ion wi h h ee
di e en endopep idases. Pep ide iden i ica ions we e ob ained by MaxQuan using a six-
ame “s op- o-s op” ansla ion o he e e ence genome sequence. Only high-quali y pep ide
iden i ica ions we e accep ed. Wi h each app oach a conside able numbe o unique p o eins
we e iden i ied. Clea ly, only a combina ion o di e en app oaches may acili a e he iden i i-
ca ion o he en i e se o small p o eins exp essed in a de ined bac e ium [15,19,49,60]. In e -
es ingly, using he 1D gel o p o ein ac iona ion, small p o eins ha e been de ec ed in
almos all ac ions indica ing s ong in e ac ions wi h la ge p o eins. In addi ion, we es ed
he alue o using di e en endop o eases o iden i ica ion o SP100: ypsin, Lys-C and
AspN. I became clea ha AspN pe o ming hyd olysis o pep ide bonds a he amine si e o
aspa yl esidues [69] was less e icien o iden i ica ion o SP100, a leas in S.au eus New-
man. No ably, in Bacillus sub ilis, Lys-C and A g-C iden i ied addi ional small p o eins no
iden i iable wi h a yp ic diges [49]. Pep ides gene a ed by AspN a e mo e equen ly acidic.
Wo hwhile emphasizing is he ac ha he pI o 26% o he iden i ied pep ides using AspN
was below ou as compa ed o only 9% o he iden i ied pep ides using Lys-C and ypsin.
The e is s ong eason o belie e ha he numbe o basic pep ides s ongly impac iden i ica-
ion o small p o eins ha end o be mo e alkaline. Fo an imp o ed iden i ica ion o basic
p o eins i may hus be essen ial o conside p o ocols aiming a highe pe cen ages o basic
pep ides.
Ou combined genome-wide p o eogenomics app oach enabled us o iden i y 175 coding
sOFRs o which a conside able numbe has no ye been desc ibed. The majo i y (n = 120) was
de ec ed by mul iple pep ides p o iding s ong e idence o hei exis ence (Fig 5A). The iden-
i ica ion o ano he 55 SP100 elied on single high-quali y PSMs. 34 ou o hese pep ides we e
iden i ied in mo e han one expe imen al app oach and 21 by a leas 10 MS/MS scans in a
leas one expe imen al app oach (S5 Table). Fo 135 o he he e iden i ied sORF RPFs ha e
been de ec ed by ibosome p o iling sugges ing ansla ional ac i i y a he espec i e genomic
egion. This echnique, howe e , p o ides only a snapsho o ansla ional ac i i y when using
a limi ed numbe o sampling poin s, as is he case he e, and i was hus no expec ed o iden-
i y ansla ional ac i i y o all sORFs iden i ied by ou MS-based app oach. In addi ion,
ansla ional ac i i y o sORFs embedded in la ge ORFs a he same s and a e ha dly o dis-
inguish om ha o he espec i e la ge ORFs by con en ional ibosomal p o iling ech-
niques. Consequen ly, ibosomal p o iling p o ided addi ional hin s o he exis ence o
sORFs iden i ied by ou p o eogenomics wo k low, howe e , u he echniques such as an i-
body based echniques o spec al lib a y-based compa isons wi h syn he ic e e ence pep ides
ha e o be included in ollow-up s udies o alida e ha in pa icula hose SP100 ha ha e
been iden i ied he e wi h only one unique pep ide a e bona ide small p o eins.
Ou o he 175 iden i ied SP100, 24 we e no co e ed by he used gene anno a ion. Th ee o
hem we e iden i ied by a leas wo unique pep ides and mo e han 50% (n = 14) we e sup-
po ed by a leas wo expe imen al app oaches. To e alua e he sui abili y o he algo i hm
applied by Peppe o deduce sORFs on he basis o iden i ied pep ides, we used a genome
anno a ion o S.au eus Newman en i ely lacking ORFs wi h up o 303 bp and he e idence
iles p o ided by MaxQuan o ou MS based app oach. In his way, 169 sORFs wi h up o
303 bp we e deduced by Peppe o which 122 o e lapped comple ely wi h sORFs anno a ed
o S.au eus Newman by NCBI (NC_009641.1; genome anno a ion om 2020-02-17). In e -
es ingly, o 24 sORFs, di e ences o sORFs anno a ed by NCBI ha e been obse ed wi h
espec o hei leng h (S6 Table). While he end o hese ORFs is clea ly de ined by one o he
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 20 / 26
h ee possible s op codons, he p edic ed s a si es a y when mul iple s a si es a e possible
and selec ion o one o hese s a si es depends on he p edic ing algo i hm used. Fo selec ion
o he bes ORF, ou analysis ool Peppe pe o ms an ORF anking based on he numbe o
iden i ied pep ides, he na u e o he s a codon, he p esence o a ibosomal binding si e and
he space be ween bo h (S6 Table). As a o emen ioned he leng h o he ORF is only ele an
when he e a e wo ORFs belonging o he same ORF class. No ably, 21 o hese sORFs
deduced by Peppe a e p eceded by a ibosomal binding si e while only se en o he sORF a -
ian s anno a ed by NCBI a e cha ac e ized by his ea u e. Th ee o hem a ied by only a ew
codons aising he ques ion as o whe he mul iple s a si es may be possible. Fo sORFSa-
New0121 we go expe imen al e idence ha ansla ion s a s a posi ion 1976846 using TTG
as s a codon. This esul s in a 104 aa p o ein ins ead o he 96 aa p o ein anno a ed by NCBI
(NWMN_RS10090). Addi ional in o ma ion p o ided by ibosome p o iling and MS based
iden i ica ion o N- e minal pep ides a e hus essen ial o imp o e ORF p edic ion [24,55,70–
72].
The biochemical oles o mos o he iden i ied SP100 a e ye o be de e mined as hey s ill
emained la gely uncha ac e ized a he molecula le el. Co-localiza ion o hei encoding
genes wi h al eady cha ac e ized genes, p esence o unc ional domains and/o subcellula
localiza ion o he p o eins in he cell [9] may p o ide i s hin s abou a possible ole o hese
p o eins in cell´s physiology and i ulence. No ably, 24 o he iden i ied SP100 a e associa ed
wi h he ou p ophage egions. Ou o hese 22 a e encoded by φNM1, φNM2 o φNM4,
which a e membe s o he Sipho i idae amily and highly simila . Ano he wo a e encoded by
φNM3 widely dis ibu ed among he human S.au eus isola es. Because o high sequence simi-
la i ies o φNM1, φNM2 o φNM4, 19 o he iden i ied p o eins a e o hologues and we e asso-
cia ed o eigh p o ein amilies.
Some sORFs a e p oximal in sequence space o ups eam o downs eam genes encoding
well cha ac e ized la ge p o eins wi h enzyma ic ac i i y such as aminopep idase
(NWMN_RS104440), glu amin amido ans e ase (NWMN_RS10105), glycosyl ans e ase
(NWMN_RS05075), ibonuclease J (NWMN_RS05355), ca diolipin syn hase
(NWMN_RS06950), and GTP- and ATP-binding p o eins (NWMN_RS02350,
NWMN_RS01520, NWMN_RS10370) in ol ed in heme biosyn hesis o DNA eplica ion.
Hence, i is easonable o hypo hesize ha he physiological ac i i y o he newly iden i ied
small p o eins migh be associa ed wi h he ac i i y o hose la ge p o eins. Mo eo e , we
iden i ied i e SP100 simila o cold shock p o eins and ou p o eins belonging o oxin-an i-
oxin sys ems. Mos in e es ingly, almos hal o he iden i ied SP100 a e basic implying a ole
in binding o mo e acidic cellula s uc u es such as nucleic acids o phospholipids. Simila
obse a ions ha e been epo ed ecen ly o small p o eins iden i ied in a simpli ied human
gu mic obiome [19]. Fu u e wo k will ocus on cha ac e izing hese p o eins.
Suppo ing in o ma ion
S1 Table. All pu a i e ORFs a e classi ied using di e en c i e ia (i wo ORF a ian s
sha e he same class, he longe ORF is p e e ed).
(DOCX)
S2 Table. Selec ed ou pu in o ma ion p o ided by Peppe .
(PDF)
S3 Table. Iden i ica ion o SP100 using di e en expe imen al wo k lows.
(XLSX)
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 21 / 26
S4 Table. Iden i ied SP100 in S aphylococcus au eus Newman.
(XLSX)
S5 Table. MS/MS coun s o SP100 iden i ied by one unique pep ide.
(XLSX)
S6 Table. Anno a ed sORFs di e en ly p edic ed by Peppe .
(XLSX)
S1 MS MS Spec a. High-quali y MS/MS spec a o yp ic pep ides unique o he iden i-
ied SP100.
(7Z)
Acknowledgmen s
We hank B. Jung o echnical assis ance and Benjamin Heinige (Ag oscope) o help wi h
imp o ing S4 Table.
Au ho Con ibu ions
Concep ualiza ion: S ephan Fuchs, Ch is ian H. Ah ens, Zoya Igna o a, Susanne Engelmann.
Da a cu a ion: S ephan Fuchs, Ma in Kucklick.
Fo mal analysis: Ma in Kucklick, E ik Lehmann.
Funding acquisi ion: S ephan Fuchs, Zoya Igna o a, Susanne Engelmann.
In es iga ion: Ma in Kucklick, E ik Lehmann, Alexande Beckmann, Maya Wilkens, Baban
Kol e, Ay en Mus a aye a, Tobias Ludwig, Mau ice Diwo.
Me hodology: Ma in Kucklick, Jose Wissing.
P ojec adminis a ion: S ephan Fuchs, Zoya Igna o a, Susanne Engelmann.
Resou ces: Lo ha Ja¨nsch, Susanne Engelmann.
So wa e: S ephan Fuchs, Alexande Beckmann.
Supe ision: S ephan Fuchs, Zoya Igna o a, Susanne Engelmann.
W i ing – o iginal d a : S ephan Fuchs, Susanne Engelmann.
W i ing – e iew & edi ing: S ephan Fuchs, Ma in Kucklick, E ik Lehmann, Alexande Beck-
mann, Maya Wilkens, Baban Kol e, Ay en Mus a aye a, Tobias Ludwig, Mau ice Diwo,
Jose Wissing, Lo ha Ja¨nsch, Ch is ian H. Ah ens, Zoya Igna o a, Susanne Engelmann.
Re e ences
1. Lowy FD. S aphylococcus au eus in ec ions. N Engl J Med. 1998; 339: 520–32. h ps://doi.o g/10.1056/
NEJM199808203390806 PMID: 9709046
2. Te elin H, Riley D, Ca u o C, Medini D. Compa a i e genomics: he bac e ial pan-genome. Cu Opin
Mic obiol. 2008; 11: 472–7. h ps://doi.o g/10.1016/j.mib.2008.09.006 PMID: 19086349
3. Bosi E, Monk JM, Aziz RK, Fondi M, Nize V, Palsson BO. Compa a i e genome-scale modelling o
S aphylococcus au eus s ains iden i ies s ain-speci ic me abolic capabili ies linked o pa hogenici y.
P oc Na l Acad Sci U S A. 2016; 113: E3801–9. h ps://doi.o g/10.1073/pnas.1523199113 PMID:
27286824
4. Kusch H, Engelmann S. Sec e s o he sec e ome in S aphylococcus au eus. In J Med Mic obiol. 2014;
304: 133–41. h ps://doi.o g/10.1016/j.ijmm.2013.11.005 PMID: 24424242
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 22 / 26
5. Beche D, Hempel K, Sie e s S, Zu¨hlke D, Pane
´-Fa e
´J, O o A, e al. A p o eomic iew o an impo an
human pa hogen— owa ds he quan i ica ion o he en i e S aphylococcus au eus p o eome. PLoS
One. 2009; 4: e8176. h ps://doi.o g/10.1371/jou nal.pone.0008176 PMID: 19997597
6. Fuchs S, Zu¨hlke D, Pane
´-Fa e
´J, Kusch H, Wol C, Reiss S, e al. Au eolib—a p o eome signa u e
lib a y: owa ds an unde s anding o S aphylococcus au eus pa hophysiology. PLoS One. 2013; 8:
e70669. h ps://doi.o g/10.1371/jou nal.pone.0070669 PMID: 23967085
7. Zu¨hlke D, Do
¨ ies K, Be nha d J, Maass S, Mun el J, Liebsche V, e al. Cos s o li e—Dynamics o he
p o ein in en o y o S aphylococcus au eus du ing anae obiosis. Sci Rep. 2016; 6: 28172. h ps://doi.
o g/10.1038/s ep28172 PMID: 27344979
8. Zieband AK, Kusch H, Degne M, Jagli z S, Sibbald MJ, A ends JP, e al. P o eomics unco e s ex eme
he e ogenei y in he S aphylococcus au eus exop o eome due o genomic plas ici y and a ian gene
egula ion. P o eomics. 2010; 10: 1634–44. h ps://doi.o g/10.1002/pmic.200900313 PMID: 20186749
9. S ekho en DJ, Omasi s U, Queba e M, Dehio C, Ah ens CH. P o eome-wide iden i ica ion o p edomi-
nan subcellula p o ein localiza ions in a bac e ial model o ganism. J P o eomics. 2014; 99: 123–37.
h ps://doi.o g/10.1016/j.jp o .2014.01.015 PMID: 24486812
10. Lipman DJ, Sou o o A, Koonin EV, Panchenko AR, Ta uso a TA. The ela ionship o p o ein conse -
a ion and sequence leng h. BMC E ol Biol. 2002; 2: 20. h ps://doi.o g/10.1186/1471-2148-2-20
PMID: 12410938
11. Rudd KE, Humphe y-Smi h I, Wasinge VC, Bai och A. Low molecula weigh p o eins: a challenge o
pos -genomic esea ch. Elec opho esis. 1998; 19: 536–44. h ps://doi.o g/10.1002/elps.1150190413
PMID: 9588799
12. Hemm MR, Paul BJ, Schneide TD, S o z G, Rudd KE. Small memb ane p o eins ound by compa a i e
genomics and ibosome binding si e models. Mol Mic obiol. 2008; 70: 1487–501. h ps://doi.o g/10.
1111/j.1365-2958.2008.06495.x PMID: 19121005
13. Dinge ME, Pang KC, Me ce TR, Ma ick JS. Di e en ia ing p o ein-coding and noncoding RNA: chal-
lenges and ambigui ies. PLoS Compu Biol. 2008; 4: e1000176. h ps://doi.o g/10.1371/jou nal.pcbi.
1000176 PMID: 19043537
14. C appe J, Van C iekinge W, T ooskens G, Hayakawa E, Luy en W, Bagge man G, e al. Combining in
silico p edic ion and ibosome p o iling in a genome-wide sea ch o no el pu a i ely coding sORFs.
BMC Genomics. 2013; 14: 648. h ps://doi.o g/10.1186/1471-2164-14-648 PMID: 24059539
15. Ma J, Died ich JK, Jung eis I, Donaldson C, Vaughan J, Kellis M, e al. Imp o ed Iden i ica ion and Anal-
ysis o Small Open Reading F ame Encoded Polypep ides. Anal Chem. 2016; 88: 3967–75. h ps://doi.
o g/10.1021/acs.analchem.6b00191 PMID: 27010111
16. Mi a e -Ve de S, Fe a T, Espadas-Ga cia G, Mazzolini R, Gha ab A, Sabido E, e al. Un a eling he
hidden uni e se o small p o eins in bac e ial genomes. Mol Sys Biol. 2019; 15: e8290. h ps://doi.o g/
10.15252/msb.20188290 PMID: 30796087
17. Omasi s U, Va ada ajan AR, Schmid M, Goe ze S, Melidis D, Bou qui M, e al. An in eg a i e s a egy o
iden i y he en i e p o ein coding po en ial o p oka yo ic genomes by p o eogenomics. Genome Res.
2017; 27: 2083–95. h ps://doi.o g/10.1101/g .218255.116 PMID: 29141959
18. Pauli A, Valen E, Schie AF. Iden i ying (non-)coding RNAs and small pep ides: challenges and oppo u-
ni ies. Bioessays. 2015; 37: 103–12. h ps://doi.o g/10.1002/bies.201400103 PMID: 25345765
19. Pe uschke H, Ande s J, S adle PF, Jehmlich N, on Be gen M. En ichmen and iden i ica ion o small
p o eins in a simpli ied human gu mic obiome. J P o eomics. 2020; 213: 103604. h ps://doi.o g/10.
1016/j.jp o .2019.103604 PMID: 31841667
20. Sbe o H, F emin BJ, Zli ni S, Ed o s F, G een ield N, Snyde MP, e al. La ge-Scale Analyses o
Human Mic obiomes Re eal Thousands o Small, No el Genes. Cell. 2019; 178: 1245–59 e14. h ps://
doi.o g/10.1016/j.cell.2019.07.016 PMID: 31402174
21. Sla o SA, Mi chell AJ, Schwaid AG, Cabili MN, Ma J, Le in JZ, e al. Pep idomic disco e y o sho
open eading ame-encoded pep ides in human cells. Na Chem Biol. 2013; 9: 59–64. h ps://doi.o g/
10.1038/nchembio.1120 PMID: 23160002
22. Yang X, Tschaplinski TJ, Hu s GB, Jawdy S, Ab aham PE, Lank o d PK, e al. Disco e y and anno a-
ion o small p o eins using genomics, p o eomics, and compu a ional app oaches. Genome Res. 2011;
21: 634–41. h ps://doi.o g/10.1101/g .109280.110 PMID: 21367939
23. Yang X, Jensen SI, Wul T, Ha ison SJ, Long KS. Iden i ica ion and alida ion o no el small p o eins
in Pseudomonas pu ida. En i on Mic obiol Rep. 2016; 8: 966–74. h ps://doi.o g/10.1111/1758-2229.
12473 PMID: 27717237
24. Wea e J, Mohammad F, Buski k AR, S o z G. Iden i ying Small P o eins by Ribosome P o iling wi h
S alled Ini ia ion Complexes. mBio. 2019; 10: e02819–18. h ps://doi.o g/10.1128/mBio.02819-18
PMID: 30837344
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 23 / 26
25. Washie l S, Findeiss S, Mulle SA, Kalkho S, on Be gen M, Ho acke IL, e al. RNAcode: obus dis-
c imina ion o coding and noncoding egions in compa a i e sequence da a. RNA. 2011; 17: 578–94.
h ps://doi.o g/10.1261/ na.2536111 PMID: 21357752
26. Yeasmin F, Yada T, Akimi su N. Mic opep ides Encoded in T ansc ip s P e iously Iden i ied as Long
Noncoding RNAs: A New Chap e in T ansc ip omics and P o eomics. F on Gene . 2018; 9: 144.
h ps://doi.o g/10.3389/ gene.2018.00144 PMID: 29922328
27. B ylinski M. Explo ing he "da k ma e " o a mammalian p o eome by p o ein s uc u e and unc ion
modeling. P o eome Sci. 2013; 11: 47. h ps://doi.o g/10.1186/1477-5956-11-47 PMID: 24321360
28. O MW, Mao Y, S o z G, Qian SB. Al e na i e ORFs and small ORFs: shedding ligh on he da k p o e-
ome. Nucleic Acids Res. 2020; 48: 1029–42. h ps://doi.o g/10.1093/na /gkz734 PMID: 31504789
29. S o z G, Wol YI, Ramamu hi KS. Small p o eins can no longe be igno ed. Annu Re Biochem. 2014;
83: 753–77. h ps://doi.o g/10.1146/annu e -biochem-070611-102400 PMID: 24606146
30. Wang F, Xiao J, Pan L, Yang M, Zhang G, Jin S, e al. A sys ema ic su ey o mini-p o eins in bac e ia
and a chaea. PLoS One. 2008; 3: e4027. h ps://doi.o g/10.1371/jou nal.pone.0004027 PMID:
19107199
31. Wang R, B augh on KR, K e schme D, Bach TH, Queck SY, Li M, e al. Iden i ica ion o no el cy oly ic
pep ides as key i ulence de e minan s o communi y-associa ed MRSA. Na u e Med. 2007; 13: 1510–
4. h ps://doi.o g/10.1038/nm1656 PMID: 17994102
32. Be nheime AW, Rudy B. In e ac ions be ween memb anes and cy oly ic pep ides. Biochimica Biophy-
sica Ac a. 1986; 864: 123–41. h ps://doi.o g/10.1016/0304-4157(86)90018-3 PMID: 2424507
33. Peschel A, O o M. Phenol-soluble modulins and s aphylococcal in ec ion. Na u e Re Mic obiol. 2013;
11: 667–73. h ps://doi.o g/10.1038/n mic o3110 PMID: 24018382
34. Ve don J, Gi a din N, Lacombe C, Be jeaud JM, Hecha d Y. del a-hemolysin, an upda e on a mem-
b ane-in e ac ing pep ide. Pep ides. 2009; 30: 817–23. h ps://doi.o g/10.1016/j.pep ides.2008.12.017
PMID: 19150639
35. Du hie ES, Lo enz LL. S aphylococcal coagulase; mode o ac ion and an igenici y. J Gen Mic obiol.
1952; 6: 95–107. h ps://doi.o g/10.1099/00221287-6-1-2-95 PMID: 14927856
36. B ussow H, Canchaya C, Ha d WD. Phages and he e olu ion o bac e ial pa hogens: om genomic
ea angemen s o lysogenic con e sion. Mic obiol Mol Biol Re . 2004; 68: 560–602. h ps://doi.o g/10.
1128/MMBR.68.3.560-602.2004 PMID: 15353570
37. Gill SR, Fou s DE, A che GL, Mongodin EF, Deboy RT, Ra el J, e al. Insigh s on e olu ion o i ulence
and esis ance om he comple e genome analysis o an ea ly me hicillin- esis an S aphylococcus
au eus s ain and a bio ilm-p oducing me hicillin- esis an S aphylococcus epide midis s ain. J Bac e -
iol. 2005; 187: 2426–38. h ps://doi.o g/10.1128/JB.187.7.2426-2438.2005 PMID: 15774886
38. Diep BA, Gill SR, Chang RF, Phan TH, Chen JH, Da idson MG, e al. Comple e genome sequence o
USA300, an epidemic clone o communi y-acqui ed me icillin- esis an S aphylococcus au eus. Lance .
2006; 367: 731–9. h ps://doi.o g/10.1016/S0140-6736(06)68231-7 PMID: 16517273
39. Bae T, Baba T, Hi ama su K, Schneewind O. P ophages o S aphylococcus au eus Newman and hei
con ibu ion o i ulence. Mol Mic obiol. 2006; 62: 1035–47. h ps://doi.o g/10.1111/j.1365-2958.2006.
05441.x PMID: 17078814
40. Baba T, Bae T, Schneewind O, Takeuchi F, Hi ama su K. Genome sequence o S aphylococcus au eus
s ain Newman and compa a i e analysis o s aphylococcal genomes: polymo phism and e olu ion o
wo majo pa hogenici y islands. J Bac e iol. 2008; 190: 300–10. h ps://doi.o g/10.1128/JB.01000-07
PMID: 17951380
41. Reiss S, Pane
´-Fa e
´J, Fuchs S, F anc¸ois P, Liebeke M, Sch enzel J, e al. Global analysis o he S aph-
ylococcus au eus esponse o mupi ocin. An imic Agen s Chemo he . 2012; 56: 787–804. h ps://doi.
o g/10.1128/AAC.05363-11 PMID: 22106209
42. Laemmli UK. Clea age o s uc u al p o eins du ing he assembly o he head o bac e iophage T4.
Na u e. 1970; 227: 680–5. h ps://doi.o g/10.1038/227680a0 PMID: 5432063
43. Le ch MF, Schoen elde SMK, Ma incola G, Wencke FDR, Ecka M, Fo
¨ s ne KU, e al. A non-coding
RNA om he in e cellula adhesion (ica) locus o S aphylococcus epide midis con ols polysaccha ide
in e cellula adhesion (PIA)-media ed bio ilm o ma ion. Mol Mic obiol. 2019; 111: 1571–91. h ps://doi.
o g/10.1111/mmi.14238 PMID: 30873665
44. Kumme A, Nishan h G, Koschel J, Klawonn F, Schlu¨ e D, Ja
¨nsch L. Lis e iosis down egula es hepa ic
cy och ome P450 enzymes in suble hal mu ine in ec ion. P o eomics Clin Appl. 2016; 10: 1025–35.
h ps://doi.o g/10.1002/p ca.201600030 PMID: 27273978
45. Buli a B, Zusch a e W, Be nal I, B ude D, Klawonn F, on Be gen M, e al. P o eomic de ini ion o
human mucosal-associa ed in a ian T cells de e mines hei unique molecula e ec o pheno ype. Eu
J Immunol. 2018; 48: 1336–49. h ps://doi.o g/10.1002/eji.201747398 PMID: 29749611
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 24 / 26
46. Del Campo C, Ba holoma
¨us A, Fedyunin I, Igna o a Z. Seconda y S uc u e ac oss he Bac e ial T an-
sc ip ome Re eals Ve sa ile Roles in mRNA Regula ion and Func ion. PLoS Gene . 2015; 11:
e1005613. h ps://doi.o g/10.1371/jou nal.pgen.1005613 PMID: 26495981
47. Ba holoma
¨us A, Kol e B, Mus a aye a A, Goebel I, Fuchs S, Benndo E, Engelmann S, e al. smOR-
Fe : a modula algo i hm o de ec small ORFs in p oka yo es. Nucl Acids Res doi:gkab477
48. Cassidy L, P asse D, Linke D, Schmi z RA, Tholey A. Combina ion o Bo om-up 2D-LC-MS and Semi-
op-down GelF ee-LC-MS Enhances Co e age o P o eome and Low Molecula Weigh Sho Open
Reading F ame Encoded Pep ides o he A chaeon Me hanosa cina mazei. J P o eome Res. 2016; 15:
3773–83. h ps://doi.o g/10.1021/acs.jp o eome.6b00569 PMID: 27557128
49. Ba el J, Va ada ajan AR, Su a T, Ah ens CH, Maass S, Beche D. Op imized P o eomics Wo k low o
he De ec ion o Small P o eins. J P o eome Res. 2020; 19: 4004–18. h ps://doi.o g/10.1021/acs.
jp o eome.0c00286 PMID: 32812434
50. Swaney DL, Wenge CD, Coon JJ. Value o using mul iple p o eases o la ge-scale mass spec ome-
y-based p o eomics. J P o eome Res. 2010; 9: 1323–9. h ps://doi.o g/10.1021/p 900863u PMID:
20113005
51. Yuan P, D’Lima NG, Sla o SA. Compa a i e Memb ane P o eomics Re eals a Nonanno a ed E.coli
Hea Shock P o ein. Biochemis y. 2018; 57: 56–60. h ps://doi.o g/10.1021/acs.biochem.7b00864
PMID: 29039649
52. Yin X, Wu O M, Wang H, Hobbs EC, Shabalina SA, S o z G. The small p o ein Mg S and small RNA
Mg R modula e he Pi A phospha e sympo e o boos in acellula magnesium le els. Mol Mic obiol.
2019; 111: 131–44. h ps://doi.o g/10.1111/mmi.14143 PMID: 30276893
53. Fon aine F, Fuchs RT, S o z G. Memb ane localiza ion o small p o eins in Esche ichia coli. J Biol
Chem. 2011; 286: 32464–74. h ps://doi.o g/10.1074/jbc.M111.245696 PMID: 21778229
54. Zhou M, Boekho s J, F ancke C, Siezen RJ. Loca eP: genome-scale subcellula -loca ion p edic o o
bac e ial p o eins. BMC Bioin o ma ics. 2008; 9: 173. h ps://doi.o g/10.1186/1471-2105-9-173 PMID:
18371216
55. Meydan S, Ma ks J, Klepacki D, Sha ma V, Ba ano PV, Fi h AE, e al. Re apamulin-Assis ed Ribo-
some P o iling Re eals he Al e na i e Bac e ial P o eome. Mol Cell. 2019; 74: 481–93 e6. h ps://doi.
o g/10.1016/j.molcel.2019.02.017 PMID: 30904393
56. Impens F, Rolhion N, Radoshe ich L, Beca in C, Du al M, Mellin J, e al. N- e minomics iden i ies
P li42 as a memb ane minip o ein conse ed in Fi micu es and c i ical o s essosome ac i a ion in Lis-
e ia monocy ogenes. Na Mic obiol. 2017; 2: 17005. h ps://doi.o g/10.1038/nmic obiol.2017.5 PMID:
28191904
57. Ingolia NT, Ghaemmaghami S, Newman JR, Weissman JS. Genome-wide analysis in i o o ansla-
ion wi h nucleo ide esolu ion using ibosome p o iling. Science. 2009; 324: 218–23. h ps://doi.o g/10.
1126/science.1168978 PMID: 19213877
58. Go ochowski TE, Chelyshe a I, E iksen M, Nai P, Pede sen S, Igna o a Z. Absolu e quan i ica ion o
ansla ional egula ion and bu den using combined sequencing app oaches. Mol Sys Biol. 2019; 15:
e8719. h ps://doi.o g/10.15252/msb.20188719 PMID: 31053575
59. Ven u ini E, S ensson SL, Maaß S, Gelhausen R, Eggenho e F, Li L, e al. A global da a-d i en census
o Salmonella small p o eins and hei po en ial unc ions in bac e ial i ulence. mic oLi e. 2020;1.
60. Pe uschke H, Scho i C, Canzle S, Riesbeck S, Poehlein A, Daniel R, e al. Disco e y o no el commu-
ni y- ele an small p o eins in a simpli ied human in es inal mic obiome. Mic obiome. 2021; 9: 55.
h ps://doi.o g/10.1186/s40168-020-00981-z PMID: 33622394
61. D’Lima NG, Khi un A, Rosenbloom AD, Yuan P, Gassaway BM, Ba be KW, e al. Compa a i e P o eo-
mics Enables Iden i ica ion o Nonanno a ed Cold Shock P o eins in E.coli. J P o eome Res. 2017; 16:
3722–31. h ps://doi.o g/10.1021/acs.jp o eome.7b00419 PMID: 28861998
62. Bazzini AA, Johns one TG, Ch is iano R, Mackowiak SD, Obe maye B, Fleming ES, e al. Iden i ica ion
o small ORFs in e eb a es using ibosome oo p in ing and e olu iona y conse a ion. EMBO J.
2014; 33: 981–93. h ps://doi.o g/10.1002/embj.201488411 PMID: 24705786
63. Cassidy L, Helbig AO, Kaulich PT, Weidenbach K, Schmi z RA, Tholey A. Mul idimensional sepa a ion
schemes enhance he iden i ica ion and molecula cha ac e iza ion o low molecula weigh p o eomes
and sho open eading ame-encoded pep ides in op-down p o eomics. J P o eomics. 2021; 230:
103988. h ps://doi.o g/10.1016/j.jp o .2020.103988 PMID: 32949814
64. Omasi s U, Ah ens CH, Mulle S, Wollscheid B. P o e : in e ac i e p o ein ea u e isualiza ion and in e-
g a ion wi h expe imen al p o eomic da a. Bioin o ma ics. 2014; 30: 884–6. h ps://doi.o g/10.1093/
bioin o ma ics/b 607 PMID: 24162465
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 25 / 26