scieee Science in your language
[en] (orig)

Towards the characterization of the hidden world of small proteins in Staphylococcus aureus, a proteogenomics approach.

Abstract

Small proteins play essential roles in bacterial physiology and virulence, however, automated algorithms for genome annotation are often not yet able to accurately predict the corresponding genes. The accuracy and reliability of genome annotations, particularly for small open reading frames (sORFs), can be significantly improved by integrating protein evidence from experimental approaches. Here we present a highly optimized and flexible bioinformatics workflow for bacterial proteogenomics covering all steps from (i) generation of protein databases, (ii) database searches and (iii) peptide-to-genome mapping to (iv) visualization of results. We used the workflow to identify high quality peptide spectrum matches (PSMs) for small proteins (≤ 100 aa, SP100) in Staphylococcus aureus Newman. Protein extracts from S. aureus were subjected to different experimental workflows for protein digestion and prefractionation and measured with highly sensitive mass spectrometers. In total, 175 proteins with up to 100 aa (SP100) were identified. Out of these 24 (ranging from 9 to 99 aa) were novel and not contained in the used genome annotation.144 SP100 are highly conserved and were found in at least 50% of the publicly available S. aureus genomes, while 127 are additionally conserved in other staphylococci. Almost half of the identified SP100 were basic, suggesting a role in binding to more acidic molecules such as nucleic acids or phospholipids.

Read accessible full text

Towards the characterization of the hidden world of small proteins in Staphylococcus aureus, a proteogenomics approach.

Author: Fuchs, Stephan,Kucklick, Martin,Lehmann, Erik,Beckmann, Alexander,Wilkens, Maya,Kolte, Baban,Mustafayeva, Ayten,Ludwig, Tobias,Diwo, Maurice,Wissing, Josef,Jänsch, Lothar,Ahrens, Christian H,Ignatova, Zoya,Engelmann, Susanne
Publisher: PLOS
Year: 2021
DOI: 10.1371/journal.pgen.1009585
Source: https://repository.helmholtz-hzi.de/bitstream/10033/622939/1/Fuchs%20et%20al.pdf
RESEARCH ARTICLE
Towa ds he cha ac e iza ion o he hidden
wo ld o small p o eins in S aphylococcus
au eus, a p o eogenomics app oach
S ephan Fuchs
1
, Ma in KucklickID
2,3
, E ik LehmannID
2,3
, Alexande BeckmannID
2,3
,
Maya WilkensID
1,2,3
, Baban Kol e
4
, Ay en Mus a aye a
2,3
, Tobias Ludwig
2,3
,
Mau ice DiwoID
2,3
, Jose Wissing
5
, Lo ha Ja
¨nsch
5
, Ch is ian H. Ah ensID
6
,
Zoya Igna o aID
4
, Susanne EngelmannID
2,3
*
1Robe Koch Ins i u e, Me hodenen wicklung und Fo schungsin as uk u (MF), Be lin, Ge many,
2Uni e si y o Technical Sciences B aunschweig, Ins i u e o Mic obiology, B aunschweig, Ge many,
3Helmhol z Cen e o In ec ion Resea ch GmbH, Mic obial P o eomics, B aunschweig, Ge many,
4Uni e si y o Hambu g, Ins i u e o Biochemis y and Molecula Biology, Hambu g, Ge many, 5Helmhol z
Cen e o In ec ion Resea ch GmbH, Cellula P o eomics, B aunschweig, Ge many, 6Ag oscope, Resea ch
G oup Molecula Diagnos ics, Genomics and Bioin o ma ics & SIB Swiss Ins i u e o Bioin o ma ics, Basel,
Swi ze land
*Susanne.Engelm[email p o ec ed]
Abs ac
Small p o eins play essen ial oles in bac e ial physiology and i ulence, howe e , au o-
ma ed algo i hms o genome anno a ion a e o en no ye able o accu a ely p edic he co -
esponding genes. The accu acy and eliabili y o genome anno a ions, pa icula ly o small
open eading ames (sORFs), can be signi ican ly imp o ed by in eg a ing p o ein e idence
om expe imen al app oaches. He e we p esen a highly op imized and lexible bioin o ma -
ics wo k low o bac e ial p o eogenomics co e ing all s eps om (i) gene a ion o p o ein
da abases, (ii) da abase sea ches and (iii) pep ide- o-genome mapping o (i ) isualiza ion
o esul s. We used he wo k low o iden i y high quali y pep ide spec um ma ches (PSMs)
o small p o eins (�100 aa, SP100) in S aphylococcus au eus Newman. P o ein ex ac s
om S.au eus we e subjec ed o di e en expe imen al wo k lows o p o ein diges ion and
p e ac iona ion and measu ed wi h highly sensi i e mass spec ome e s. In o al, 175 p o-
eins wi h up o 100 aa (SP100) we e iden i ied. Ou o hese 24 ( anging om 9 o 99 aa)
we e no el and no con ained in he used genome anno a ion.144 SP100 a e highly con-
se ed and we e ound in a leas 50% o he publicly a ailable S.au eus genomes, while
127 a e addi ionally conse ed in o he s aphylococci. Almos hal o he iden i ied SP100
we e basic, sugges ing a ole in binding o mo e acidic molecules such as nucleic acids o
phospholipids.
Au ho summa y
Con en ional au oma ic genome anno a ion algo i hms o en neglec open eading
ames smalle han 300 nucleo ides (sORF). The e a e se e al easons hinde ing
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 1 / 26
a1111111111
a1111111111
a1111111111
a1111111111
a1111111111
OPEN ACCESS
Ci a ion: Fuchs S, Kucklick M, Lehmann E,
Beckmann A, Wilkens M, Kol e B, e al. (2021)
Towa ds he cha ac e iza ion o he hidden wo ld o
small p o eins in S aphylococcus au eus, a
p o eogenomics app oach. PLoS Gene 17(6):
e1009585. h ps://doi.o g/10.1371/jou nal.
pgen.1009585
Edi o : Kai Papen o , F ied ich-Schille -Uni e si a
Jena, GERMANY
Recei ed: No embe 24, 2020
Accep ed: May 7, 2021
Published: June 1, 2021
Pee Re iew His o y: PLOS ecognizes he
bene i s o anspa ency in he pee e iew
p ocess; he e o e, we enable he publica ion o
all o he con en o pee e iew and au ho
esponses alongside inal, published a icles. The
edi o ial his o y o his a icle is a ailable he e:
h ps://doi.o g/10.1371/jou nal.pgen.1009585
Copy igh : ©2021 Fuchs e al. This is an open
access a icle dis ibu ed unde he e ms o he
C ea i e Commons A ibu ion License, which
pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided he o iginal
au ho and sou ce a e c edi ed.
Da a A ailabili y S a emen : The mass
spec ome y p o eomics da a ha e been deposi ed
o he P o eomeXchange Conso ium (h p://
au oma ic anno a ion and p edic ion o sho genes: (i) sORFs possess insu icien
sequence in o ma ion o domain and homology sea ch, (ii) only a limi ed numbe o
expe imen ally alida ed sORFs can se e as empla es, and (iii) sORFs show he endency
o be species-speci ic. We hus es ablished a p o eogenomics wo k low, which is execu ed
by wo open sou ce ools, Sal and Peppe (h ps://gi lab.com/s. uchs/peppe ), and uses
pep ide da a ob ained by mass spec ome y o iden i ica ion o genes in bac e ia ha a e
ha dly p edic able by au oma ic anno a ion algo i hms. As a p oo o concep , we selec ed
S aphylococcus au eus, one o he mos equen ly sequenced bac e ia and iden i ied 36
p o eins no ye conside ed in he used genome anno a ion o S.au eus Newman. 24 he e
o a e no el small p o eins wi h up o 100 aa (SP100) in S.au eus Newman. This clea ly
demons a es ha ou wo k low is ideally sui ed o imp o e gene anno a ion o al eady
anno a ed bac e ial genomes. In he u u e, i may also acili a e p o ein and ORF de ec-
ion in no anno a ed bac e ial genomes.
In oduc ion
S aphylococcus au eus is a G am-posi i e human pa hogen o g ea clinical impo ance. S.
au eus causes mainly nosocomial in ec ions in immunocomp omized pa ien s, which a e e-
quen ly associa ed wi h di icul o ea mul id ug- esis an S.au eus pheno ypes [1]. Wi h
11,809 genome sequences (including 576 comple e genomes), which a e publicly a ailable in
he e e ence sequence da abase o he Na ional Cen e o Bio echnology In o ma ion (Re Seq;
s a us 2020-08-19), S.au eus is among he mos equen ly sequenced bac e ia. The numbe o
anno a ed open eading ames anges om 2,411 o 3,147 pe comple e genome sequence.
The en i e pan-genome o S.au eus has no ye been desc ibed, due o he ac ha he geno-
mic di e si y o S.au eus is e y high [2,3]. Howe e , a p elimina y S.au eus pan-genome
based on he compa ison o 64 S.au eus genome sequences is composed o 7,411 genes, o
which abou 20% a e conse ed cons i u ing he co e-genome [3]. The highes a iabili y has
been ound among genes coding o ex acellula and su ace-associa ed p o eins [4] which is
o pa icula impo ance as hese p o eins a e essen ially in ol ed in di ec in e ac ions wi h
he hos en i onmen du ing in ec ion.
The p o ein in en o y o se e al S.au eus s ains has been desc ibed using highly sensi-
i e mass spec ome y (MS) echniques combined wi h liquid ch oma og aphy (LC) [5–7].
Fo S.au eus s ain COL, mo e han 1,700 p o eins (abou 60% o he heo e ical p o eome)
ha e been iden i ied, quan i ied and assigned o a ious subcellula localiza ions [5,7,8],
which can help o p edic unc ions o co-exp essed and/o co-localized p o eins [9]. How-
e e , one g oup o p o eins was highly unde ep esen ed in he S.au eus p o eome: e y
small p o eins no longe han 100 amino acids (aa) (= SP100). Al oge he , Beche and col-
leagues [5] de ec ed 82 anno a ed SP100 o which only ou p o eins we e below 50 aa in
leng h (= SP50).
The expe imen al de ec ion o SP100 by sho gun p o eomics is di icul and addi ionally
hampe ed by he ac ha he co esponding sho open eading ames (sORFs) a e o en
o e looked by con en ional genome anno a ion algo i hms. The e a e se e al easons hinde -
ing au oma ed p edic ion and accu a e anno a ion o sORFs, such as insu icien sequence
in o ma ion o domain and homology sea ches, a limi ed numbe o expe imen ally alida ed
empla es, and hei endency o species speci ici y [10–12]. Hence, di e en ia ion be ween
sORFs wi h low and high coding po en ial is challenging and he numbe o alse posi i es
among p edic ed sORFs is ex emely high [13]. Gi en hese ac s, genome anno a ions
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 2 / 26
p o eomecen al.p o eomexchange.o g) ia he
PRIDE pa ne eposi o y (Vizcaino JA, Cso das A,
del-To o N, Dianes JA, G iss J, La idas I, e al.
2016 upda e o he PRIDE da abase and i s ela ed
ools. Nucleic Acids Res. 2016;44: D447-56.) wi h
he da ase iden i ie PXD017932. High Quali y MS/
MS spec a o he iden i ied yp ic pep ides unique
o SP100 a e accessible in Supplemen al
Ma e ials. Ribosome p o iling da a ha e been
deposi ed wi hin Gene Exp ession Omnibus (GEO)
unde accession numbe GSE150601.
Funding: This wo k was unded by he Deu sche
Fo schungsgemeinscha (h ps://www.d g.de/)
(GRK PROCOMPAS) o SE, by he Deu sche
Fo schungsgemeinscha (GRK PROCOMPAS) o
LJ, by he Deu sche Fo schungsgemeinscha
(INST 188/365-1 FUGG DFG) o SE, by he
Schweize ische Na ional onds zu Fo¨ de ung de
Wissenscha lichen Fo schung (h p://www.sn .ch/
de/Sei en/de aul .aspx) (197391) o CHA, by he
Deu sche Fo schungsgemeinscha (IG 73/16-1
SPP 2002) o ZI. The unde s had no ole in s udy
design, da a collec ion and analysis, decision o
publish, o p epa a ion o he manusc ip .
Compe ing in e es s: The au ho s ha e decla ed
ha no compe ing in e es s exis .
ou inely used a bi a y cu -o s o a minimum ORF leng h o 50 o 100 codons. In addi ion,
he low molecula weigh o hese p o eins complica es expe imen al isola ion and educes he
numbe o MS-compa ible pep ides.
O e he las yea s, a ious a emp s ha e been made ha add ess one o bo h o hese
issues. This includes expe imen al app oaches such as ibosome p o iling and p o eogenomics
o iden i y his g oup o p o eins as well as bioin o ma ics app oaches o a mo e eliable p e-
dic ion and comp ehensi e anno a ion o sORFs [14–26]. Fo ins ance, di e en compu a-
ional app oaches ha e been de eloped o sORF p edic ion, which ha e in common ha he
coding po en ial o a pu a i e sORF is sco ed based on one o mo e ea u es such as nucleo ide
composi ion, synonymous and non-synonymous subs i u ion a es, phylogene ic conse a ion
o p o ein domain de ec ion [13,16].
Despi e hese majo challenges, he e is no doub ha small p o eins play a pi o al ole in
essen ial cellula p ocesses, hence, i is ex emely impo an o imp o e ou abili y o unco e
his pool o hidden p o eins [20,27–30]. Fi s unc ional cha ac e iza ions p o e hei in ol e-
men in a ious cellula p ocesses such as p o ein olding, egula ion o gene exp ession, mem-
b ane anspo , p o ein modi ica ion and signal ansduc ion in di e en bac e ia ( o an
o e iew see [30]). In addi ion, some small p o eins ha e an ex acellula unc ion and exhibi
oxic o an imic obial ac i i y. In e es ingly, mos o he small p o eins cha ac e ized so a a e
associa ed wi h he cell memb ane and a e poo ly conse ed a he sequence le el [29]. In S.
au eus, he mos p ominen small p o eins a e phenol soluble modulins wi h a leng h o 20 o
40 aa [31] and del a-hemolysin (26 aa) [32]. Phenol soluble modulins possess mul iple oles in
S.au eus pa hogenesis by inducing cell lysis o blood cells, s imula ing in lamma o y esponses
and in luencing bio ilm o ma ion ( o e iew see [33]). Del a-hemolysin in e ac s wi h mem-
b anes o a ious blood cells, which concen a ion dependen ly esul s in a memb ane dis u -
bance and e en in cell lysis ( o e iew see [34]). While bo h phenol soluble modulins and
del a-hemolysin ha e been s udied in de ail in ecen yea s, da a on he iden i ica ion and
unc ional cha ac e iza ion o o he SP50 in S.au eus a e almos comple ely missing.
The a ailabili y o nume ous S.au eus genome sequences de ines i as a well-sui ed model
bac e ium o p edic ion and iden i ica ion o small p o eins and pep ides. The numbe o
anno a ed coding sequences wi h up o 303 nucleo ides is highly a iable in he 576 comple e
S.au eus genome sequences anging om 287 o 621 SP100. This is mainly a ibu ed o he
ac ha a ious algo i hms o genome anno a ion we e applied. Among S.au eus e e ence
s ains, s ain Newman plays a pi o al ole. Fi s isola ed in 1952 om a human in ec ion [35],
i is one o he mos equen ly used S.au eus s ains in in ec ion models as i is cha ac e ized
by a ela i ely s able pheno ype. In addi ion, ou p ophages we e iden i ied in he genome o
s ain Newman, inse ed a di e en si es in he ch omosome, exceeding he egula ly
obse ed numbe o p ophages in S.au eus [36–38]. In a mu ine in ec ion model, he loss o
all ou p ophages signi ican ly educed he i ulence po en ial o he s ain [39]. The exis ence
o hese p ophages made S.au eus Newman an excellen model o s udy hei impac on i u-
lence and cells physiology. The genome sequence o s ain Newman, was p edic ed o encode
a leas 2,854 p o eins, again, he numbe o anno a ed SP100 is a he low and amoun s o 493
p o eins [40] (NC_009641.1; genome anno a ion om 2020-02-17).
To mo e comp ehensi ely iden i y p o eins wi h up o 100 aa, we used S.au eus Newman
as a model sys em and de eloped an ully- ea u ed p o eogenomics wo k low ha combines in
silico ansla ion o he en i e genome sequence, a ious LC-MS/MS wo k lows, and a bioin-
o ma ics pipeline o pep idomics da a analyses. The wo k low is highly op imized, lexible
and eady o use in o he bac e ial species.
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 3 / 26
Me hods
Bac e ial s ains, cul i a ion condi ions and cell lysis
S.au eus Newman [35] was cul i a ed in 100 mL complex medium (TSB) a 37˚C and 120
pm o an op ical densi y a 540 nm (OD
540
) o 1 and 7. Cells we e ha es ed by mul iple cen-
i uga ion s eps and dis up ed by cell homogeniza ion (Fas P ep-24, MP Biomedicals) ( o
de ails see [41]). All expe imen s ha e been pe o med wi h h ee biological eplica es. The
p o ein concen a ion was de e mined using he Ro i-Nanoquan assay (Ro h, Ka ls uhe, Ge -
many) and he p o ein solu ion was s o ed a -20˚C.
F ac iona ion o p o eins and pep ides and p o eoly ic clea age
Gel-based app oach. 40 μg o cy oplasmic p o eins we e sepa a ed by one dimensional
SDS polyac ylamide gel elec opho ese (1D SDS PAGE) acco ding o Laemmli [42] wi h he
ollowing modi ica ions: he loading bu e consis ed o 3.75% ( / ) glyce ol, 1.25% ( / ) ß-
me cap oe hanol, 0.6% (w/ ) SDS, 0.0014% (w/ ) b omophenol blue, 16.5 mM T is-HCl (pH
6.8). The sepa a ion gel con ained 12% (w/ ) ac ylamide gel (wi h 0.32% bisac ylamide), 0.375
M T is-HCL (pH 8.8), 0.255% (w/ ) SDS, 0.062% (w/ ) APS, and 0.062% ( / ) TEMED and
he s acking gel 5% (w/ ) ac ylamide (wi h 0.13% (w/ ) bisac ylamide), 0.125 M T is-HCl (pH
6.8) 0.25% (w/ ) SDS, 0.075% (w/ ) APS, and 0.075% ( / ) TEMED.
P o eins we e ixed wi h 40% ( / ) e hanol and 10% ( / ) ace ic acid o one hou and sub-
sequen ly s ained wi h colloidal coomassie [43] o one hou . In-gel diges ion using ypsin
and ex ac ion o he pep ides we e ca ied ou as desc ibed by Le ch e al. [43] wi h an addi-
ional ex ac ion s ep using ace oni ile.
Fo diges ion wi h Lys-C, a bu e con aining 25 mM TRIS/HCl and 1 mM EDTA (pH 8.5)
was used. The applied enzyme concen a ion was 1/40 o he o al p o ein concen a ion.
Diges ion o AspN was pe o med in 10 mM T is-HCl (pH 8.0) wi h a inal AspN concen a-
ion o 1/50 o he o al p o ein concen a ion.
Gel- ee app oach. The gel- ee app oach was pe o med by applying yp ic in-solu ion
diges ion ollowed by an Oasis HLB-SPE-ca idge pu i ica ion and SCX- ac iona ion. In
de ail, 40 μg o c ude p o ein ex ac we e sol ed in 8 M u ea and 2 M hiou ea and adjus ed o
a inal concen a ion o 6 M u ea. A e addi ion o 1.6 μL o 5 mM DTT in 50 mM ammonium
bica bona e solu ion (pH 7.8), he p o ein solu ion was incuba ed o 30 min a oom empe a-
u e. Fo alkyla ion, 1 μL o a eshly p epa ed 55 mM IAA in 50 mM ammonium bica bona e
bu e was added o 10 μL o he p o ein solu ion and incuba ed o 20 min in he da k a
oom empe a u e. Subsequen ly, he solu ion was adjus ed o a inal concen a ion o 1 mM
CaCl
2
and 1 M u ea using CaCl
2
sol ed in a 50 mM ammonium bica bona e bu e . Fo diges-
ion, 1 μg ypsin (in 50 mM ammonium bica bona e and 1 mM CaCl
2
) was applied o 50 μg
p o ein. Diges ion was pe o med o 12 h a 37˚C wi h gen le agi a ion (50 pm) and s opped
by acidi ica ion o a pH alue o �2.5 wi h 10% o mic acid.
Fo pep ide pu i ica ion, Oasis HLB-SPE-ca idges (1cc, 10mg, Wa e s, Mil o d, MA,
USA) we e ini ially condi ioned wi h ace oni ile and hen wi h 0.5% o mic acid in 60% ace o-
ni ile. Subsequen ly, ca idges we e equilib a ed wi h wo olumes o 0.5% o mic acid. Sam-
ples we e loaded on he ca idges; he low- h ough was collec ed and again loaded on he
ca idge. The pep ides we e washed i e imes wi h 0.5% o mic acid (FA) and elu ed wice
wi h 0.85 mL 60% ACN 0.5% FA. Elua es we e d ied in a speed ac (Eppendo Concen a o
plus, Eppendo AG, Hambu g, Ge many) and ozen a -20˚C.
SCX ac iona ion was done as desc ibed by Kumme e al. [44]. To educe he numbe o
ac ions o eigh , pep ide-con aining ac ions we e combined.
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 4 / 26
Pep ide desal ing. ZipTips (C18, Me ck Millipo e, Bille ica, MA, USA) we e condi ioned
wi h 50% ace oni ile wice. Subsequen ly hey we e equilib a ed h ee imes wi h 0.1% FA in
5% ace oni ile. 10 μL o each pep ide ac ion ( esol ed in 20 μL 0.1% FA in 5% ace oni ile
o 60 min) we e loaded on he C18 ma ix o he ip by aspi a ing 10 imes. Elu ion was pe -
o med h ee imes by aspi a ing i e imes wi h 0.1% FA in 60% ace oni ile in a new mic o
es ube. Samples we e d ied in a speed ac.
Liquid ch oma og aphy coupled mass spec ome y (LC-MS/MS)
Fo LC-MS/MS, each pep ide ac ion o a sample was sol ed in 16 μL o 0.1% FA in 3% ace o-
ni ile o one hou , ul asonica ed in a wa e ba h o 5 min and ul acen i uged.
O bi ap Velos P o MS. LC-MS/MS uns wi h he O bi ap Velos P o MS (The mo
Fishe Scien i ic Inc, Wal ham, MA USA) we e done as desc ibed by Le ch and cowo ke s
[43].
O bi ap Fusion MS. LC-MS sys em and used columns a e desc ibed by Buli a and
cowo ke s [45]. A 200 min g adien was applied, s a ing wi h 3.7% bu e B (80% ace oni ile,
5% DMSO and 0.1% o mic acid) and 96.3% bu e A (0.1% o mic acid, 5% DMSO): 0–5 min
3.7% B; 5–125 min 3.7–31.3% B; 125–165 min 31.3–62.5% B; 165–172 min 62.5–90.0% B; 172–
177 min 90% B; 177–182 min 90–3.7% B, 182–200 min 3.7% B.
P ima y Scans we e pe o med a he O bi ap in he p o ile modus scanning an m/z o
350–1800 wi h a esolu ion ( ull wid h a hal maximum a m/z 400) o 120,000 and a lock
mass o 445.1200. Using he Xcalibu so wa e (The mo Fishe Scien i ic Inc., San Jose, CA,
USA), he mass spec ome e was con olled and ope a ed in he “ op speed” mode, allowing
he au oma ic selec ion o as much as possible wice o ou old-cha ged pep ides in a h ee-
second ime window, and he subsequen agmen a ion o hese pep ides. In he non- a ge ed
modus, p ima y ions (±10 ppm) we e selec ed by he quad upole (isola ion window: 1.6 m/z),
agmen ed in he ion ap using a da a dependen CID mode ( op speed mode, 3 seconds) o
he mos abundan p ecu so ions wi h an exclusion ime o 13 s and analysed by he ion ap.
P o ein da abase gene a ion
To conside i s ull coding po en ial, he genome sequence o he S.au eus subsp. au eus s ain
Newman (NC_009641.1) was ansla ed in all six eading ames om s op o s op codon
using he SALT ool (h ps://gi lab.com/s. uchs/peppe ). Genome ci cula i y has been consid-
e ed and en ies wi h less han 9 amino acids excluded esul ing in a o al o 177,532 sequence
en ies.
MS Da a analysis and s a is ics
Analyses o he ob ained MS and MS/MS da a we e pe o med using MaxQuan (Max Planck
Ins i u e o Biochemis y, Ma ins ied, Ge many, www.maxquan .o g, e sion 1.5.2.8) and he
ollowing pa ame e s: pep ide ole ance 5 ppm; a ole ance o agmen ions o 0.6 Da; a i-
able modi ica ions: me hionine oxida ion and ace yla ion a p o ein N- e minus, ixed modi i-
ca ion: ca bamidome hyla ion (Cys); a maximum o wo missed clea ages and ou
modi ica ions pe pep ide was allowed. Fo he iden i ica ion o SP100, a minimum o one
unique pep ide pe p o ein and a ixed alse disco e y a e (FDR) o 0.0001 o PSMs and 0.01
o p o eins was applied. The minimum sco e was se o 40 o unmodi ied and modi ied pep-
ides, he minimum del a sco e was se o 6 o unmodi ied pep ides and o 17 o modi ied
pep ides. All samples we e sea ched agains he S.au eus T ansla ion Da abase (TRDB) wi h a
decoy mode o e e ed sequences and common con aminan s supplied by MaxQuan
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 5 / 26

Fo iden i ica ion o non-anno a ed open eading ames based on iden i ied pep ides a
p o eogenomics ool has been de eloped ha desc ibes any iden i ied pep ide in he degene -
a ed DNA code sequence. By his, exac ma ches wi hin he e e ence genome can be ound
and, subsequen ly, il e ed based on exis ing anno a ion and loca ion.
Phylogene ic and unc ional analyses
Fo phylogene ic analyses we downloaded all comple e genome sequences o S.au eus
(n = 541) and s aphylococci (n = 165) om NCBI Re Seq (s a e o 2020-05-18). All SP100
sequences we e sea ched agains he downloaded genome sequences using blas n. Based on
he bes hi alignmen o e e y bac e ial ch omosome, he iden i y ela ed o he ull que y
leng h was calcula ed. Only alignmen s sha ing a leas 90% iden i y wi h he ull que y
sequence we e conside ed. Based on his, ela i e species and genus conse a ion a es ha e
been calcula ed.
Fo unc ional analyses we sea ched all SP100 sequences agains he eggNOG da abase 5.0
using eggNOG-mappe 2.0 (de aul pa ame e s). The axonomic scope was au oma ically
adjus ed o each que y o ensu e co ec classi ica ion o phage p o eins. Only unc ions om
one- o-one o hology we e ans e ed.
Ribosome p o iling
Lib a y p epa a ion. S.au eus cells om 30 mL cul u e g own in TSB medium o OD
550
= 1 we e ha es ed by apid cen i uga ion and esuspended in 390 μL ice cold 20 mM T is
lysis bu e pH 8.8, con aining 10 mM MgCl
2
x 6 H
2
O, 100 mM NH
4
Cl, 20 mM T is (pH 8.0),
0.4% T i on-X-100, 4 U DNase, 0.4 μL Supe ase-In (Ambion), 1mM chlo amphenicol. Cells
we e dis up ed by cell homogeniza ion (Fas P ep-24, MP Biomedicals) wi h 0.5 mL glass
beads (diame e 0.1 mm) o 30 s a 6.5 m/s ollowed by incuba ion on ice o 5 min. These
s eps we e epea ed wice. To emo e cell deb is, cell lysa es we e cen i uged and subsequen ly
s o ed a -80˚C and 100 A
260
uni s o ibosome-bound mRNA ac ion we e subjec ed o
nucleoly ic diges ion wi h 10 uni s/μl mic ococcal nuclease (The mo ishe ) in bu e wi h
pH 9.2 (10 mM T is pH 11 con aining 50 mM NH
4
Cl, 10 mM MgCl
2
, 0.2% i on X-100,
100 μg/mL chlo amphenicol and 20 mM CaCl
2
). The RNA agmen s we e deple ed using
S.au eus iboPOOL RNA oligo se (siTOOLs, Ge many) and he lib a y p epa a ion was pe -
o med as p e iously desc ibed [46].
Bioin o ma ic analyses o ibosome p o iling RNAs. Raw sequencing eads we e
immed using FASTX Toolki (quali y h eshold: 20) and adap e s we e cu using cu adap
(minimal o e lap o 1 n ) and mapped o he genome e sion NC_009641.1 (NCBI, Janua y
2020). Following ex ac ion o eads mapping o RNAs, he emaining eads we e uniquely
mapped o he e e ence genome using Bow ie, pa ame e se ings: -l 16 -n 1 -e 50 -m 1—
s a a–bes y. Non-uniquely mapped eads we e non-conside ed ( o mo e de ails see [47]).
Resul s
Es ima ing he numbe o spu ious sORFs in S.au eus Newman using
single-nucleo ide pe mu a ion es ing
A single nucleo ide pe mu a ion es was used o e i y he global con idence and signi icance
o ORFs based on hei leng h only. Fo his pu pose, ORFs we e de ec ed in he genome
sequence o S.au eus Newman (NC_009641.1; NCBI ansla ion able 11; longes ORFs p e-
e ed). As alse-posi i e es ima e, we used he median numbe o ORFs de ec ed in pe mu ed
genome sequences (n = 1000), ha show he same nucleo ide composi ion bu in andom
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 6 / 26
o de and can he e o e be assumed o no con ain any biological in o ma ion (Fig 1A).
Acco dingly, he alse disco e y a e (FDR) o ORFs de ec ed in he biological sequence is
113% o hose wi h a maximum leng h o 63 bp (coding o 20 aa), 133% o hose wi h a
leng h be ween 66 and 153 bp (21 o 50 aa) and 94% o hose wi h a leng h be ween156 and
303 bp (51 o 100 aa). This highligh s he need o addi ional e idence o he eliable anno a-
ion o sORFs. In con as , coding sequences (CDS) wi h a leng h o a leas 396 bp (�131 aa)
did no occu by chance in he pe mu ed sequence se . Since s a and s op codons end o be
AT- ich, he FDR o sORFs inc eases wi h an inc easing GC con en o an o ganism (Fig 1B).
C ea ing mo e comp ehensi e p o ein da abases o S.au eus Newman
using in silico ansla ion
To iden i y small p o eins no co e ed in he Re Seq anno a ion, we gene a ed a p o ein da a-
base conside ing he ull coding po en ial o S.au eus Newman by ansla ing all six eading
ames o he espec i e genome sequence and c ea ing a sepa a e p o ein en y o each
sequence be ween wo s op codons. The esul ing TRansla ion Da aBase (TRDB) comp ises
177,532 sequences wi h a minimum leng h o 9 aa. We es ablished an au oma ed wo k low o
he ansla ion o (ci cula ) bac e ial genome sequences (Fig 2A). The co esponding py hon-
based ool called Sal is publicly a ailable (h ps://gi lab.com/s. uchs/peppe ). I allows he
ex ac ion o all po en ial ORFs om a ci cula o linea genome sequence and suppo s bo h
s a o s op and s op o s op codon ex ac ion (acco ding o NCBI ansla ion able 11). The
ex ac ed sequences can be au oma ically ansla ed and, i equi ed, diges ed in silico in o
indi idual pep ides using p ede ined diges ion pa e ns o di e en enzymes. All DNA, p o ein
and pep ide sequences can be s o ed in indi idual FASTA iles. The espec i e sequence heade
can be ully cus omized o mee speci ic equi emen s. Mo eo e , o each sequence collec ion,
ea u e ables can be expo ed as abula o delimi ed ex iles lis ing di e en physicochemical
Fig 1. sORF equencies in biological and pe mu ed genome sequences. (A) Es ima ion o he p opo ion o alse-posi i e sORFs in p edic ions based solely on
s a and s op codons: All po en ial ORF sequences (NCBI ansla ion able 11; longes ORF a ian s p e e ed) we e ex ac ed om he genome sequence o S.
au eus Newman and 1,000 pe mu ed sequence de i a i es showing he same nucleo ide composi ion bu in andom o de and can he e o e be assumed o no longe
con ain any biological in o ma ion. Resul ing ORFs we e binned based on hei leng h. Bin sizes a e shown o he genuine e e ence sequence (o ange) and he
pe mu ed sequences (g ey; as median) up o a maximum ORF leng h o 303 bp (= 100 aa). Especially small ORFs end o occu andomly. (B) Impac o GC con en
on he numbe o spu ious ORF: Acco ding o (A), bin sizes a e gi en o genome sequences and hei pe mu ed sequence de i a i es (n = 1,000; as median) wi h
a ying GC con en . The used e e ence genome sequences we e NC_000913.3 (Esche ichia coli K-12 subs . MG1655; 50.8%GC; 4,641,652 bp), NC_000964.3
(Bacillus sub ilis subsp.sub ilis s . 168; 43.5%GC; 4,215,606 bp), NC_007633.1 (Mycoplasma cap icolum subsp.cap icolum ATCC 27343; 23.8%GC, 1,010,023 bp),
NC_009641.1 (S aphylococcus au eus subsp.au eus s . Newman; 32.9%GC; 2,878,897 bp), NC_010162.1 (So angium cellulosum So ce56; 71.4%GC; 1,303,779 bp).
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g001
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 7 / 26
Fig 2. Bac e ial p o eogenomics wo k low p o ided by Sal and Peppe . (A) C ea ion o p o ein and pep ide
da abases using Sal : Based on a FASTA ile as inpu , Sal ex ac s all po en ial ORFs om a gi en (ci cula ) genome
sequence using di e en me hods (s op o s op codon, s a o s op codon). The esul ing ORF sequences a e hen
ansla ed in silico in o p o ein sequences ha can be u he diges ed using di e en in silico p o eases (T yspin,
Chymo ypsin, Asp-N, Lys-C and P o einase K). Fo each le el (ORFs, p o eins, pep ides) indi idual FASTA iles and
ables ( ab-sepa a ed alues, TSV) a e c ea ed, lis ing a ious sequence-de i ed p ope ies such as molecula weigh o
isoelec ic poin s. (B) P o eogenomics analyses using Peppe : Pep ide o spec um ma ches (PSMs) ob ained om
di e en samples a e di ec ly ex ac ed om MaxQuan e idence iles (MQ). Spec al da a can be ex ensi ely e alua ed
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 8 / 26
p ope ies such as molecula weigh s, isoelec ic poin s o g and a e age o hyd opa hy
(GRAVY) alues. To he bes o ou knowledge, his is he only eely a ailable ool ha o e s
such a a ie y o unc ions ( u he in o ma ion see h ps://gi lab.com/s. uchs/peppe ).
C ea ing a ully au oma ed ye lexible wo k low o bac e ial
p o eogenomics
To deduce pu a i e ORFs om a lis o iden i ied pep ides, we c ea ed a ule-based expe sys-
em called Peppe which is ully au oma ed and op imized o bac e ial p o eogenomics (Fig
2B). In b ie , sequences and quali y measu es o pep ide spec um ma ches (PSMs) a e
ex ac ed om e idence iles (and, op ionally, ms/ms iles) p o ided by he MaxQuan so -
wa e (Max Planck Ins i u e o Biochemis y, Ma ins ied, Ge many, e sion 1.5.2.8; h p://
www.maxquan .o g). Di e en spec um- and quali y-based il e c i e ia can be hen applied
au oma ically o es ic he analyses o high-quali y PSMs only (a leas 5 consecu i e y- o b-
ions o a leas 2 �4 y-ions o b-ions o a leas 4 b- and 4 y-ions) (see also Table 1). The espec-
i e il e c i e ia we e deduced om a e y ex ensi e isual inspec ion and assessmen o he
MS/MS spec a by expe s. The objec i e was o educe he numbe o alse-posi i es when
applying a cu -o o only one unique pep ide pe p o ein o iden i ica ion o SP100. We
equi ed ha he e be a sequence ag o a leas i e consecu i e b o y agmen ions o wo
imes ou consecu i e b o y ions in a spec um o be conside ed [21]. In addi ion, pep ide
speci ic mass acks should ha e signi ican le els abo e he backg ound le els. Hence, Peppe
is able o au oma ically sco e and il e high-quali y PSMs on hese speci ic equi emen s and
o apply addi ional il e s ela ed o he in ensi y co e age (>0.1), And omeda Sco e (�40)
and pos e io e o p obabili y (<0.1) (Table 1). High-quali y MS/MS spec a o yp ic pep-
ides unique o SP100 in S.au eus a e accessible in he Supplemen al Ma e ials.
In he nex s ep, all po en ial coding si es (DNA ma ches) a e iden i ied o each high-qual-
i y PSM wi hin he gi en genome sequence conside ing he degene a ed na u e o he DNA
code. On he basis o he DNA ma ches ound, po en ial ORFs a e deduced acco ding o he
ollowing ules: i s , an ORF mus con ain all successi e DNA ma ches ha a e encoded on
he same s and and in he same eading ame and no sepa a ed by an in e posed s op
codon. Secondly, he ORF is ex ended un il he i s s op codon downs eam o he las DNA
ma ch and encoded in he same eading ame. In he inal s ep, p edic ing he ansla ional
s a si e, he mos ups eam DNA ma ch co e ed by he po en ial ORF plays an essen ial ole.
Th ee di e en cases can be dis inguished he e: (1) The iden i ied pep ide encoded by he
mos ups eam DNA ma ch is no a p o eoly ic (e.g. yp ic) p oduc . In his case, he i s
codon o he mos ups eam DNA ma ch is assumed o be he ansla ion s a si e. Since N-
e minal me hionine esidues a e clea ed om a numbe o bac e ial p o eins du ing
o apply di e en spec um quali y and eplica ion c i e ia can be de ined o es ic he analysis o highly- eliable PSMs
only. Respec i e coding si es a e de e mined in a gi en (ci cula ) genome sequence p o ided as FASTA ile. The
esul ing coding si es (DNA ma ches), ha can be il e ed e.g. by exclusi i y a e used o p edic he pu a i e open
eading ames. Addi ional in o ma ion such as po en ial ibosomal binding si es, gene syn eny based on he e e ence
genome anno a ion, and conse a ion in gi en sequence collec ions (p o ided as FASTA iles) a e collec ed. Di e en
iles a e c ea ed o a chi e all analysis pa ame e s (log ile), esul s on pep ide, DNA ma ch, and ORF le el (TSV), and
an upda ed e e ence genome anno a ion in eg a ing he iden i ied DNA ma ches and ORFs (Genbank ile; GB). (C)
Resul s isualiza ion: GB iles c ea ed by Peppe can be used o esul s isualiza ion using hi d-pa y so wa e (he e:
Geneious P ime, Bioma e s L d.). The genome sequence (black line) wi h coo dina es is shown on op. Exis ing
anno a ions a e highligh ed in yellow and g een. ORFs wi h he highes coding po en ial ega ding Peppe and po en ial
ibosomal binding si es (RBS) a e highligh ed in ed. DNA ma ches a e show in ligh - ed, i he espec i e pep ide is no
encoded elsewhe e in he genome (exclusi e ma ch), o ligh -blue, i mul iple coding si es exis o he espec i e
pep ide.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g002
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 9 / 26
me hionine. Fo he emaining 14 sORFs, any possible in- ame s a codon was assigned and
he esul ing ORFs we e e alua ed based on di e en c i e ia implemen by Peppe (S1 Table).
While he majo i y o he 24 new sORFs has been sugges ed o s a wi h ATG (n = 11), six
sORFs p esumably ini ia e a non-canonical ansla ional s a si es such as TTG (n = 2), ATT
(n = 1), GTG (n = 1), ATA (n = 1) and ATC (n = 1). Fo se en sORFs, ini ia ion a comple ely
unexpec ed codons was pos ula ed. In hese cases, he iden i ied pep ides encoded by he mos
ups eam DNA ma ch o he espec i e ORF was no a p o eoly ic p oduc and he i s codon
o he mos ups eam DNA ma ch was assumed o be he ansla ion s a si e (Fig 9A). To
ind u he suppo o he p edic ed sORFs, we looked o possible ups eam ibosomal bind-
ing si es. Hence, eigh o he newly de i ed sORFs a e p eceded by a ibosomal binding si e
wi hin a dis ance up o 14 nucleo ides ups eam o he pu a i e s a codon.
Based on hei genome localiza ion, he 24 no el sORFs can be classi ied in o h ee main
ca ego ies: (i) loca ed in in e genic egions, (ii) o e lapping wi h o he ORFs ei he wi h he 5´
o wi h he 3´-end bu in a di e en eading ame, and (iii) loca ed wi hin anno a ed p o ein
coding sequence bu in a di e en eading ame. Fo he la e wo we can addi ionally dis in-
guish be ween hose loca ed a he same s and and hose loca ed a he complemen a y s and.
The majo i y (n = 13) belongs o he g oup (iii) o which i e a e localized a he same s and
(Fig 9B). This g oup o p o eins is highly in e es ing and i s e idence o hei exis ence was
ecen ly epo ed o se e al o he o ganisms using N- e minomics o ibosomal p o iling in
combina ion wi h e apamulin [24,28,55,56]. Se en o he newly iden i ied sORFs we e
de ec ed by a leas wo di e en app oaches. Fou sORFs we e alloca ed o g oup (ii) and en
o g oup (i). No ably, h ee SP100 belonging o g oup (i) a e encoded by egions wi hin pseu-
dogenes (Fig 10A–10C). These a e sORFSaNew0004 (pseudogene NWMN_RS15675),
sORFSaNew0010 (pseudogene NWMN_RS01305), and sORFSaNew0044 (pseudogene
NWMN_RS04585). While o NWMN_RS01305 only one ame shi mu a ion leads o an
Fig 7. Phylogene ic conse a ion o he iden i ied SP100 a species and genus le el. SP100 we e sea ched agains he
Re Seq genome sequences using blas n. Based on he bes hi alignmen o e e y genome, he iden i y ela ed o he
ull que y leng h was calcula ed. Only alignmen s sha ing 90% iden i y wi h he ull-leng h que y sequence we e
conside ed. On he basis o hese esul s, ela i e species and genus conse a ion a es ha e been calcula ed.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g007
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 16 / 26

Fig 8. Func ional classi ica ion o he iden i ied SP100. Conse ed p o ein domains we e iden i ied using he
eggNOG da abase. On he basis o his, 140 SP100 (80%) we e success ully assigned o an eggNOG o hologues clus e
p o iding addi ional suppo o hei exis ence by biological signi icance.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g008
Fig 9. Cha ac e is ics o no ye anno a ed SP100. In o al 24 SP100 we e iden i ied in S.au eus Newman, which
we e no co e ed by he used gene anno a ion (NC 009641.1: genome anno a ion om 2020-02-17). The encoding
sORFs we e de i ed by Peppe on he basis o he iden i ied pep ides and speci ic c i e ia conce ning he ansla ional
s a codon, he Shine Dalga no sequence and he leng h o he space be ween bo h. (A) Dis ibu ion o di e en
ansla ional s a codons be ween he newly p edic ed sORFs. (B) Cha ac e is ics o he genome localiza ion o he
newly p edic ed sORFs: (i) in e genic egions, (ii) pa ly o e lapping wi h ano he ORF a he same s and o a he
complemen a y s and, and (iii) comple ely o e lapping wi h ano he ORF a he same s and o a he complemen a y
s and.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g009
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 17 / 26
in e up ion o he open eading ame and ansla ion o he ull leng h p o ein canno be
excluded, o NWMN_RS15675 and NWMN_RS01305 se e al in e up ions ha e been
de ec ed and ansla ion o he ull leng h p o ein is ex emely unlikely.
T ansla ion o iden i ied sORFs
We addi ionally pe o med ibosome p o iling o S.au eus g own unde he exac same
g ow h condi ions as used o assessing he p o eome. Ribosome p o iling p o ides a snapsho
o ansla ion [57] and he posi ion o he ansla ing ibosomes can be assessed wi h codon
Fig 10. Iden i ied sORFs localized wi hin pseudogenes. Schema ic p esen a ion o he NWMN_RS15675 ( npA) (A)
NWMN_RS01305 (B) and NWMN_RS04585 locus (C) based on he anno a ion o he S.au eus Newman genome
sequence (NC_009641.1; genome anno a ion om 2020-02-17). Anno a ed pseudogenes a e shown in ligh g een and
he de i ed coding sequence (CDS) in yellow. Ma ched unique pep ides iden i ied by MS/MS a e depic ed in da k
g een and he bes ORF de i ed by Peppe on he basis o he iden i ied unique pep ides and addi ional ea u es is
depic ed in da k ed. Peppe analyses o diges ion wi h ypsin a e p esen ed.
h ps://doi.o g/10.1371/jou nal.pgen.1009585.g010
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 18 / 26
p ecision [58]. The ob ained sequencing eads, which ep esen ibosome-p o ec ed mRNA
agmen s (RPF), we e mapped o he genome o S.au eus Newman and ansla ion p o iles
we e gene a ed ( o mo e de ails see [47]). Fo 135 (76%) sORFs iden i ied by using ou p o-
eogenomics pipeline we de ec ed RFPs. Among hem a e i e ha we e missing in he used
genome anno a ion. Fo an addi ional se en o he newly anno a ed sORFs, based on ou
app oaches, eads ha e been mapped o he espec i e coding egion close o he signal o
noise h eshold; hey would equi e o he alida ion expe imen s.
In case o sORFs embedded wi hin p o ein coding egions a he same s and, we we e no
able o clea ly assign RPFs. In Esche ichia coli, Re apamulin-enhanced Ribo-seq analysis
(Ribo-RET), which combines e apamulin speci ic a es s o ini ia ing ibosomes and ibo-
some p o iling o map ansla ional s a si es, e ealed pu a i e in e nal s a si es in a numbe
o genes [24,55]. Howe e , p elimina y a emp s o use his echnique in S.au eus o mapping
ansla ional ini ia ion si es ha e ailed so a , possibly due o an ine icien anspo o e apa-
mulin in o he cell.
27 sORFs, which we e iden i ied by he p o eogenomics app oach, showed no ansla ional
ac i i y a all. Among hem a e 15 sORFs which a e localized on h ee di e en lysogenic
phages. I is in e es ing o no e, ha hese phages a e highly simila and since we only used
uniquely mapped eads (i.e. mapping o only one posi ion o he genome), i is likely ha he
RPFs we e disca ded due o ambigui y. In sum, we ha e obse ed ansla ional ac i i y o he
majo i y (76%) o he sORFs de ec ed by ou p o eogenomics app oach, alida ing he p edic-
i e powe o he used MS da a se s and TRDB in combina ion wi h he he e de eloped p o eo-
genomics ool Peppe o da a analyses. (S4 Table).
Discussion
P edic ion and iden i ica ion o sORFs coding o p o eins smalle han 100 aa (SP100) is s ill
e y challenging om bo h compu a ional and expe imen al poin s o iew. In ecen yea s,
se e al s udies s a ed o add ess his issue in a sys ema ic way by combining compu a ional
p edic ion and expe imen al alida ion using ibosome p o iling and/o mass spec ome y
[16,17,19,22,24,48,49,59–63] P o eomics based iden i ica ion o sORFs elying on mass spec-
ome y is especially ad an ageous as i p o ides di ec e idence o he exis ence o sORF
encoded p o eins and alida es no only ansla ion o sORFs bu also s abili y o hei gene
p oduc s [64,65]. Howe e , mass spec ome y based iden i ica ion o sORFs is aced by se -
e al challenges: he e y low numbe o pep ides a ailable o mass spec ome y and he
dependence on amino acid e e ence sequences. To add ess his, we he e de eloped a highly
op imized and lexible da a analysis pipeline o bac e ial p o eogenomics, co e ing all s eps
om (i) p o ein da abase gene a ion, (ii) da abase sea ch, (iii) pep ide- o-genome mapping,
and (i ) esul isualiza ion. The wo k low is based on ou bac e ial p o eogenomics pipeline
Peppe , ex ended by Sal , a genome ansla o gene a ing p o ein and pep ide da abases using
di e en me hods (e.g., s op- o-s op o s a - o-s op ansla ion). Peppe ep esen s a ule
based expe sys em ha enables ully au oma ed MS da a analysis o empi ical and e idence
based gene anno a ion. Au oma ic ule based selec ion o high quali y PSMs o pep ide iden i-
ica ion combined wi h an imp o ed s a codon de ec ion based on N- e minal pep ides dis-
inguishes his wo k low om al eady exis ing pipelines o bac e ial p o eogenomics
[16,17,66–68]. I is hus pa icula ly well placed o iden i y exp essed SP100 in bac e ia by one
unique pep ide comple ely independen o exis ing genome anno a ions. Compa ed o e y
sensi i e ibosome p o iling app oaches, which a e mo e equen ly used o sORF iden i ica-
ion, highly con iden MS based iden i ica ion o p o eins as acili a ed by his wo k low
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 19 / 26
p o ides di ec e idence o he exis ence o small p o eins and alida es no only ansla ion
o sORFs bu also s abili y o hei gene p oduc s unde he es ed condi ions.
To e alua e commonly used expe imen al app oaches o iden i ica ion o SP100 in S.
au eus, we applied a gel-based and a gel- ee LC-MS/MS app oach in combina ion wi h h ee
di e en endopep idases. Pep ide iden i ica ions we e ob ained by MaxQuan using a six-
ame “s op- o-s op” ansla ion o he e e ence genome sequence. Only high-quali y pep ide
iden i ica ions we e accep ed. Wi h each app oach a conside able numbe o unique p o eins
we e iden i ied. Clea ly, only a combina ion o di e en app oaches may acili a e he iden i i-
ca ion o he en i e se o small p o eins exp essed in a de ined bac e ium [15,19,49,60]. In e -
es ingly, using he 1D gel o p o ein ac iona ion, small p o eins ha e been de ec ed in
almos all ac ions indica ing s ong in e ac ions wi h la ge p o eins. In addi ion, we es ed
he alue o using di e en endop o eases o iden i ica ion o SP100: ypsin, Lys-C and
AspN. I became clea ha AspN pe o ming hyd olysis o pep ide bonds a he amine si e o
aspa yl esidues [69] was less e icien o iden i ica ion o SP100, a leas in S.au eus New-
man. No ably, in Bacillus sub ilis, Lys-C and A g-C iden i ied addi ional small p o eins no
iden i iable wi h a yp ic diges [49]. Pep ides gene a ed by AspN a e mo e equen ly acidic.
Wo hwhile emphasizing is he ac ha he pI o 26% o he iden i ied pep ides using AspN
was below ou as compa ed o only 9% o he iden i ied pep ides using Lys-C and ypsin.
The e is s ong eason o belie e ha he numbe o basic pep ides s ongly impac iden i ica-
ion o small p o eins ha end o be mo e alkaline. Fo an imp o ed iden i ica ion o basic
p o eins i may hus be essen ial o conside p o ocols aiming a highe pe cen ages o basic
pep ides.
Ou combined genome-wide p o eogenomics app oach enabled us o iden i y 175 coding
sOFRs o which a conside able numbe has no ye been desc ibed. The majo i y (n = 120) was
de ec ed by mul iple pep ides p o iding s ong e idence o hei exis ence (Fig 5A). The iden-
i ica ion o ano he 55 SP100 elied on single high-quali y PSMs. 34 ou o hese pep ides we e
iden i ied in mo e han one expe imen al app oach and 21 by a leas 10 MS/MS scans in a
leas one expe imen al app oach (S5 Table). Fo 135 o he he e iden i ied sORF RPFs ha e
been de ec ed by ibosome p o iling sugges ing ansla ional ac i i y a he espec i e genomic
egion. This echnique, howe e , p o ides only a snapsho o ansla ional ac i i y when using
a limi ed numbe o sampling poin s, as is he case he e, and i was hus no expec ed o iden-
i y ansla ional ac i i y o all sORFs iden i ied by ou MS-based app oach. In addi ion,
ansla ional ac i i y o sORFs embedded in la ge ORFs a he same s and a e ha dly o dis-
inguish om ha o he espec i e la ge ORFs by con en ional ibosomal p o iling ech-
niques. Consequen ly, ibosomal p o iling p o ided addi ional hin s o he exis ence o
sORFs iden i ied by ou p o eogenomics wo k low, howe e , u he echniques such as an i-
body based echniques o spec al lib a y-based compa isons wi h syn he ic e e ence pep ides
ha e o be included in ollow-up s udies o alida e ha in pa icula hose SP100 ha ha e
been iden i ied he e wi h only one unique pep ide a e bona ide small p o eins.
Ou o he 175 iden i ied SP100, 24 we e no co e ed by he used gene anno a ion. Th ee o
hem we e iden i ied by a leas wo unique pep ides and mo e han 50% (n = 14) we e sup-
po ed by a leas wo expe imen al app oaches. To e alua e he sui abili y o he algo i hm
applied by Peppe o deduce sORFs on he basis o iden i ied pep ides, we used a genome
anno a ion o S.au eus Newman en i ely lacking ORFs wi h up o 303 bp and he e idence
iles p o ided by MaxQuan o ou MS based app oach. In his way, 169 sORFs wi h up o
303 bp we e deduced by Peppe o which 122 o e lapped comple ely wi h sORFs anno a ed
o S.au eus Newman by NCBI (NC_009641.1; genome anno a ion om 2020-02-17). In e -
es ingly, o 24 sORFs, di e ences o sORFs anno a ed by NCBI ha e been obse ed wi h
espec o hei leng h (S6 Table). While he end o hese ORFs is clea ly de ined by one o he
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 20 / 26
h ee possible s op codons, he p edic ed s a si es a y when mul iple s a si es a e possible
and selec ion o one o hese s a si es depends on he p edic ing algo i hm used. Fo selec ion
o he bes ORF, ou analysis ool Peppe pe o ms an ORF anking based on he numbe o
iden i ied pep ides, he na u e o he s a codon, he p esence o a ibosomal binding si e and
he space be ween bo h (S6 Table). As a o emen ioned he leng h o he ORF is only ele an
when he e a e wo ORFs belonging o he same ORF class. No ably, 21 o hese sORFs
deduced by Peppe a e p eceded by a ibosomal binding si e while only se en o he sORF a -
ian s anno a ed by NCBI a e cha ac e ized by his ea u e. Th ee o hem a ied by only a ew
codons aising he ques ion as o whe he mul iple s a si es may be possible. Fo sORFSa-
New0121 we go expe imen al e idence ha ansla ion s a s a posi ion 1976846 using TTG
as s a codon. This esul s in a 104 aa p o ein ins ead o he 96 aa p o ein anno a ed by NCBI
(NWMN_RS10090). Addi ional in o ma ion p o ided by ibosome p o iling and MS based
iden i ica ion o N- e minal pep ides a e hus essen ial o imp o e ORF p edic ion [24,55,70–
72].
The biochemical oles o mos o he iden i ied SP100 a e ye o be de e mined as hey s ill
emained la gely uncha ac e ized a he molecula le el. Co-localiza ion o hei encoding
genes wi h al eady cha ac e ized genes, p esence o unc ional domains and/o subcellula
localiza ion o he p o eins in he cell [9] may p o ide i s hin s abou a possible ole o hese
p o eins in cell´s physiology and i ulence. No ably, 24 o he iden i ied SP100 a e associa ed
wi h he ou p ophage egions. Ou o hese 22 a e encoded by φNM1, φNM2 o φNM4,
which a e membe s o he Sipho i idae amily and highly simila . Ano he wo a e encoded by
φNM3 widely dis ibu ed among he human S.au eus isola es. Because o high sequence simi-
la i ies o φNM1, φNM2 o φNM4, 19 o he iden i ied p o eins a e o hologues and we e asso-
cia ed o eigh p o ein amilies.
Some sORFs a e p oximal in sequence space o ups eam o downs eam genes encoding
well cha ac e ized la ge p o eins wi h enzyma ic ac i i y such as aminopep idase
(NWMN_RS104440), glu amin amido ans e ase (NWMN_RS10105), glycosyl ans e ase
(NWMN_RS05075), ibonuclease J (NWMN_RS05355), ca diolipin syn hase
(NWMN_RS06950), and GTP- and ATP-binding p o eins (NWMN_RS02350,
NWMN_RS01520, NWMN_RS10370) in ol ed in heme biosyn hesis o DNA eplica ion.
Hence, i is easonable o hypo hesize ha he physiological ac i i y o he newly iden i ied
small p o eins migh be associa ed wi h he ac i i y o hose la ge p o eins. Mo eo e , we
iden i ied i e SP100 simila o cold shock p o eins and ou p o eins belonging o oxin-an i-
oxin sys ems. Mos in e es ingly, almos hal o he iden i ied SP100 a e basic implying a ole
in binding o mo e acidic cellula s uc u es such as nucleic acids o phospholipids. Simila
obse a ions ha e been epo ed ecen ly o small p o eins iden i ied in a simpli ied human
gu mic obiome [19]. Fu u e wo k will ocus on cha ac e izing hese p o eins.
Suppo ing in o ma ion
S1 Table. All pu a i e ORFs a e classi ied using di e en c i e ia (i wo ORF a ian s
sha e he same class, he longe ORF is p e e ed).
(DOCX)
S2 Table. Selec ed ou pu in o ma ion p o ided by Peppe .
(PDF)
S3 Table. Iden i ica ion o SP100 using di e en expe imen al wo k lows.
(XLSX)
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 21 / 26

S4 Table. Iden i ied SP100 in S aphylococcus au eus Newman.
(XLSX)
S5 Table. MS/MS coun s o SP100 iden i ied by one unique pep ide.
(XLSX)
S6 Table. Anno a ed sORFs di e en ly p edic ed by Peppe .
(XLSX)
S1 MS MS Spec a. High-quali y MS/MS spec a o yp ic pep ides unique o he iden i-
ied SP100.
(7Z)
Acknowledgmen s
We hank B. Jung o echnical assis ance and Benjamin Heinige (Ag oscope) o help wi h
imp o ing S4 Table.
Au ho Con ibu ions
Concep ualiza ion: S ephan Fuchs, Ch is ian H. Ah ens, Zoya Igna o a, Susanne Engelmann.
Da a cu a ion: S ephan Fuchs, Ma in Kucklick.
Fo mal analysis: Ma in Kucklick, E ik Lehmann.
Funding acquisi ion: S ephan Fuchs, Zoya Igna o a, Susanne Engelmann.
In es iga ion: Ma in Kucklick, E ik Lehmann, Alexande Beckmann, Maya Wilkens, Baban
Kol e, Ay en Mus a aye a, Tobias Ludwig, Mau ice Diwo.
Me hodology: Ma in Kucklick, Jose Wissing.
P ojec adminis a ion: S ephan Fuchs, Zoya Igna o a, Susanne Engelmann.
Resou ces: Lo ha Ja¨nsch, Susanne Engelmann.
So wa e: S ephan Fuchs, Alexande Beckmann.
Supe ision: S ephan Fuchs, Zoya Igna o a, Susanne Engelmann.
W i ing – o iginal d a : S ephan Fuchs, Susanne Engelmann.
W i ing – e iew & edi ing: S ephan Fuchs, Ma in Kucklick, E ik Lehmann, Alexande Beck-
mann, Maya Wilkens, Baban Kol e, Ay en Mus a aye a, Tobias Ludwig, Mau ice Diwo,
Jose Wissing, Lo ha Ja¨nsch, Ch is ian H. Ah ens, Zoya Igna o a, Susanne Engelmann.
Re e ences
1. Lowy FD. S aphylococcus au eus in ec ions. N Engl J Med. 1998; 339: 520–32. h ps://doi.o g/10.1056/
NEJM199808203390806 PMID: 9709046
2. Te elin H, Riley D, Ca u o C, Medini D. Compa a i e genomics: he bac e ial pan-genome. Cu Opin
Mic obiol. 2008; 11: 472–7. h ps://doi.o g/10.1016/j.mib.2008.09.006 PMID: 19086349
3. Bosi E, Monk JM, Aziz RK, Fondi M, Nize V, Palsson BO. Compa a i e genome-scale modelling o
S aphylococcus au eus s ains iden i ies s ain-speci ic me abolic capabili ies linked o pa hogenici y.
P oc Na l Acad Sci U S A. 2016; 113: E3801–9. h ps://doi.o g/10.1073/pnas.1523199113 PMID:
27286824
4. Kusch H, Engelmann S. Sec e s o he sec e ome in S aphylococcus au eus. In J Med Mic obiol. 2014;
304: 133–41. h ps://doi.o g/10.1016/j.ijmm.2013.11.005 PMID: 24424242
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 22 / 26
5. Beche D, Hempel K, Sie e s S, Zu¨hlke D, Pane
´-Fa e
´J, O o A, e al. A p o eomic iew o an impo an
human pa hogen— owa ds he quan i ica ion o he en i e S aphylococcus au eus p o eome. PLoS
One. 2009; 4: e8176. h ps://doi.o g/10.1371/jou nal.pone.0008176 PMID: 19997597
6. Fuchs S, Zu¨hlke D, Pane
´-Fa e
´J, Kusch H, Wol C, Reiss S, e al. Au eolib—a p o eome signa u e
lib a y: owa ds an unde s anding o S aphylococcus au eus pa hophysiology. PLoS One. 2013; 8:
e70669. h ps://doi.o g/10.1371/jou nal.pone.0070669 PMID: 23967085
7. Zu¨hlke D, Do
¨ ies K, Be nha d J, Maass S, Mun el J, Liebsche V, e al. Cos s o li e—Dynamics o he
p o ein in en o y o S aphylococcus au eus du ing anae obiosis. Sci Rep. 2016; 6: 28172. h ps://doi.
o g/10.1038/s ep28172 PMID: 27344979
8. Zieband AK, Kusch H, Degne M, Jagli z S, Sibbald MJ, A ends JP, e al. P o eomics unco e s ex eme
he e ogenei y in he S aphylococcus au eus exop o eome due o genomic plas ici y and a ian gene
egula ion. P o eomics. 2010; 10: 1634–44. h ps://doi.o g/10.1002/pmic.200900313 PMID: 20186749
9. S ekho en DJ, Omasi s U, Queba e M, Dehio C, Ah ens CH. P o eome-wide iden i ica ion o p edomi-
nan subcellula p o ein localiza ions in a bac e ial model o ganism. J P o eomics. 2014; 99: 123–37.
h ps://doi.o g/10.1016/j.jp o .2014.01.015 PMID: 24486812
10. Lipman DJ, Sou o o A, Koonin EV, Panchenko AR, Ta uso a TA. The ela ionship o p o ein conse -
a ion and sequence leng h. BMC E ol Biol. 2002; 2: 20. h ps://doi.o g/10.1186/1471-2148-2-20
PMID: 12410938
11. Rudd KE, Humphe y-Smi h I, Wasinge VC, Bai och A. Low molecula weigh p o eins: a challenge o
pos -genomic esea ch. Elec opho esis. 1998; 19: 536–44. h ps://doi.o g/10.1002/elps.1150190413
PMID: 9588799
12. Hemm MR, Paul BJ, Schneide TD, S o z G, Rudd KE. Small memb ane p o eins ound by compa a i e
genomics and ibosome binding si e models. Mol Mic obiol. 2008; 70: 1487–501. h ps://doi.o g/10.
1111/j.1365-2958.2008.06495.x PMID: 19121005
13. Dinge ME, Pang KC, Me ce TR, Ma ick JS. Di e en ia ing p o ein-coding and noncoding RNA: chal-
lenges and ambigui ies. PLoS Compu Biol. 2008; 4: e1000176. h ps://doi.o g/10.1371/jou nal.pcbi.
1000176 PMID: 19043537
14. C appe J, Van C iekinge W, T ooskens G, Hayakawa E, Luy en W, Bagge man G, e al. Combining in
silico p edic ion and ibosome p o iling in a genome-wide sea ch o no el pu a i ely coding sORFs.
BMC Genomics. 2013; 14: 648. h ps://doi.o g/10.1186/1471-2164-14-648 PMID: 24059539
15. Ma J, Died ich JK, Jung eis I, Donaldson C, Vaughan J, Kellis M, e al. Imp o ed Iden i ica ion and Anal-
ysis o Small Open Reading F ame Encoded Polypep ides. Anal Chem. 2016; 88: 3967–75. h ps://doi.
o g/10.1021/acs.analchem.6b00191 PMID: 27010111
16. Mi a e -Ve de S, Fe a T, Espadas-Ga cia G, Mazzolini R, Gha ab A, Sabido E, e al. Un a eling he
hidden uni e se o small p o eins in bac e ial genomes. Mol Sys Biol. 2019; 15: e8290. h ps://doi.o g/
10.15252/msb.20188290 PMID: 30796087
17. Omasi s U, Va ada ajan AR, Schmid M, Goe ze S, Melidis D, Bou qui M, e al. An in eg a i e s a egy o
iden i y he en i e p o ein coding po en ial o p oka yo ic genomes by p o eogenomics. Genome Res.
2017; 27: 2083–95. h ps://doi.o g/10.1101/g .218255.116 PMID: 29141959
18. Pauli A, Valen E, Schie AF. Iden i ying (non-)coding RNAs and small pep ides: challenges and oppo u-
ni ies. Bioessays. 2015; 37: 103–12. h ps://doi.o g/10.1002/bies.201400103 PMID: 25345765
19. Pe uschke H, Ande s J, S adle PF, Jehmlich N, on Be gen M. En ichmen and iden i ica ion o small
p o eins in a simpli ied human gu mic obiome. J P o eomics. 2020; 213: 103604. h ps://doi.o g/10.
1016/j.jp o .2019.103604 PMID: 31841667
20. Sbe o H, F emin BJ, Zli ni S, Ed o s F, G een ield N, Snyde MP, e al. La ge-Scale Analyses o
Human Mic obiomes Re eal Thousands o Small, No el Genes. Cell. 2019; 178: 1245–59 e14. h ps://
doi.o g/10.1016/j.cell.2019.07.016 PMID: 31402174
21. Sla o SA, Mi chell AJ, Schwaid AG, Cabili MN, Ma J, Le in JZ, e al. Pep idomic disco e y o sho
open eading ame-encoded pep ides in human cells. Na Chem Biol. 2013; 9: 59–64. h ps://doi.o g/
10.1038/nchembio.1120 PMID: 23160002
22. Yang X, Tschaplinski TJ, Hu s GB, Jawdy S, Ab aham PE, Lank o d PK, e al. Disco e y and anno a-
ion o small p o eins using genomics, p o eomics, and compu a ional app oaches. Genome Res. 2011;
21: 634–41. h ps://doi.o g/10.1101/g .109280.110 PMID: 21367939
23. Yang X, Jensen SI, Wul T, Ha ison SJ, Long KS. Iden i ica ion and alida ion o no el small p o eins
in Pseudomonas pu ida. En i on Mic obiol Rep. 2016; 8: 966–74. h ps://doi.o g/10.1111/1758-2229.
12473 PMID: 27717237
24. Wea e J, Mohammad F, Buski k AR, S o z G. Iden i ying Small P o eins by Ribosome P o iling wi h
S alled Ini ia ion Complexes. mBio. 2019; 10: e02819–18. h ps://doi.o g/10.1128/mBio.02819-18
PMID: 30837344
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 23 / 26
25. Washie l S, Findeiss S, Mulle SA, Kalkho S, on Be gen M, Ho acke IL, e al. RNAcode: obus dis-
c imina ion o coding and noncoding egions in compa a i e sequence da a. RNA. 2011; 17: 578–94.
h ps://doi.o g/10.1261/ na.2536111 PMID: 21357752
26. Yeasmin F, Yada T, Akimi su N. Mic opep ides Encoded in T ansc ip s P e iously Iden i ied as Long
Noncoding RNAs: A New Chap e in T ansc ip omics and P o eomics. F on Gene . 2018; 9: 144.
h ps://doi.o g/10.3389/ gene.2018.00144 PMID: 29922328
27. B ylinski M. Explo ing he "da k ma e " o a mammalian p o eome by p o ein s uc u e and unc ion
modeling. P o eome Sci. 2013; 11: 47. h ps://doi.o g/10.1186/1477-5956-11-47 PMID: 24321360
28. O MW, Mao Y, S o z G, Qian SB. Al e na i e ORFs and small ORFs: shedding ligh on he da k p o e-
ome. Nucleic Acids Res. 2020; 48: 1029–42. h ps://doi.o g/10.1093/na /gkz734 PMID: 31504789
29. S o z G, Wol YI, Ramamu hi KS. Small p o eins can no longe be igno ed. Annu Re Biochem. 2014;
83: 753–77. h ps://doi.o g/10.1146/annu e -biochem-070611-102400 PMID: 24606146
30. Wang F, Xiao J, Pan L, Yang M, Zhang G, Jin S, e al. A sys ema ic su ey o mini-p o eins in bac e ia
and a chaea. PLoS One. 2008; 3: e4027. h ps://doi.o g/10.1371/jou nal.pone.0004027 PMID:
19107199
31. Wang R, B augh on KR, K e schme D, Bach TH, Queck SY, Li M, e al. Iden i ica ion o no el cy oly ic
pep ides as key i ulence de e minan s o communi y-associa ed MRSA. Na u e Med. 2007; 13: 1510–
4. h ps://doi.o g/10.1038/nm1656 PMID: 17994102
32. Be nheime AW, Rudy B. In e ac ions be ween memb anes and cy oly ic pep ides. Biochimica Biophy-
sica Ac a. 1986; 864: 123–41. h ps://doi.o g/10.1016/0304-4157(86)90018-3 PMID: 2424507
33. Peschel A, O o M. Phenol-soluble modulins and s aphylococcal in ec ion. Na u e Re Mic obiol. 2013;
11: 667–73. h ps://doi.o g/10.1038/n mic o3110 PMID: 24018382
34. Ve don J, Gi a din N, Lacombe C, Be jeaud JM, Hecha d Y. del a-hemolysin, an upda e on a mem-
b ane-in e ac ing pep ide. Pep ides. 2009; 30: 817–23. h ps://doi.o g/10.1016/j.pep ides.2008.12.017
PMID: 19150639
35. Du hie ES, Lo enz LL. S aphylococcal coagulase; mode o ac ion and an igenici y. J Gen Mic obiol.
1952; 6: 95–107. h ps://doi.o g/10.1099/00221287-6-1-2-95 PMID: 14927856
36. B ussow H, Canchaya C, Ha d WD. Phages and he e olu ion o bac e ial pa hogens: om genomic
ea angemen s o lysogenic con e sion. Mic obiol Mol Biol Re . 2004; 68: 560–602. h ps://doi.o g/10.
1128/MMBR.68.3.560-602.2004 PMID: 15353570
37. Gill SR, Fou s DE, A che GL, Mongodin EF, Deboy RT, Ra el J, e al. Insigh s on e olu ion o i ulence
and esis ance om he comple e genome analysis o an ea ly me hicillin- esis an S aphylococcus
au eus s ain and a bio ilm-p oducing me hicillin- esis an S aphylococcus epide midis s ain. J Bac e -
iol. 2005; 187: 2426–38. h ps://doi.o g/10.1128/JB.187.7.2426-2438.2005 PMID: 15774886
38. Diep BA, Gill SR, Chang RF, Phan TH, Chen JH, Da idson MG, e al. Comple e genome sequence o
USA300, an epidemic clone o communi y-acqui ed me icillin- esis an S aphylococcus au eus. Lance .
2006; 367: 731–9. h ps://doi.o g/10.1016/S0140-6736(06)68231-7 PMID: 16517273
39. Bae T, Baba T, Hi ama su K, Schneewind O. P ophages o S aphylococcus au eus Newman and hei
con ibu ion o i ulence. Mol Mic obiol. 2006; 62: 1035–47. h ps://doi.o g/10.1111/j.1365-2958.2006.
05441.x PMID: 17078814
40. Baba T, Bae T, Schneewind O, Takeuchi F, Hi ama su K. Genome sequence o S aphylococcus au eus
s ain Newman and compa a i e analysis o s aphylococcal genomes: polymo phism and e olu ion o
wo majo pa hogenici y islands. J Bac e iol. 2008; 190: 300–10. h ps://doi.o g/10.1128/JB.01000-07
PMID: 17951380
41. Reiss S, Pane
´-Fa e
´J, Fuchs S, F anc¸ois P, Liebeke M, Sch enzel J, e al. Global analysis o he S aph-
ylococcus au eus esponse o mupi ocin. An imic Agen s Chemo he . 2012; 56: 787–804. h ps://doi.
o g/10.1128/AAC.05363-11 PMID: 22106209
42. Laemmli UK. Clea age o s uc u al p o eins du ing he assembly o he head o bac e iophage T4.
Na u e. 1970; 227: 680–5. h ps://doi.o g/10.1038/227680a0 PMID: 5432063
43. Le ch MF, Schoen elde SMK, Ma incola G, Wencke FDR, Ecka M, Fo
¨ s ne KU, e al. A non-coding
RNA om he in e cellula adhesion (ica) locus o S aphylococcus epide midis con ols polysaccha ide
in e cellula adhesion (PIA)-media ed bio ilm o ma ion. Mol Mic obiol. 2019; 111: 1571–91. h ps://doi.
o g/10.1111/mmi.14238 PMID: 30873665
44. Kumme A, Nishan h G, Koschel J, Klawonn F, Schlu¨ e D, Ja
¨nsch L. Lis e iosis down egula es hepa ic
cy och ome P450 enzymes in suble hal mu ine in ec ion. P o eomics Clin Appl. 2016; 10: 1025–35.
h ps://doi.o g/10.1002/p ca.201600030 PMID: 27273978
45. Buli a B, Zusch a e W, Be nal I, B ude D, Klawonn F, on Be gen M, e al. P o eomic de ini ion o
human mucosal-associa ed in a ian T cells de e mines hei unique molecula e ec o pheno ype. Eu
J Immunol. 2018; 48: 1336–49. h ps://doi.o g/10.1002/eji.201747398 PMID: 29749611
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 24 / 26
46. Del Campo C, Ba holoma
¨us A, Fedyunin I, Igna o a Z. Seconda y S uc u e ac oss he Bac e ial T an-
sc ip ome Re eals Ve sa ile Roles in mRNA Regula ion and Func ion. PLoS Gene . 2015; 11:
e1005613. h ps://doi.o g/10.1371/jou nal.pgen.1005613 PMID: 26495981
47. Ba holoma
¨us A, Kol e B, Mus a aye a A, Goebel I, Fuchs S, Benndo E, Engelmann S, e al. smOR-
Fe : a modula algo i hm o de ec small ORFs in p oka yo es. Nucl Acids Res doi:gkab477
48. Cassidy L, P asse D, Linke D, Schmi z RA, Tholey A. Combina ion o Bo om-up 2D-LC-MS and Semi-
op-down GelF ee-LC-MS Enhances Co e age o P o eome and Low Molecula Weigh Sho Open
Reading F ame Encoded Pep ides o he A chaeon Me hanosa cina mazei. J P o eome Res. 2016; 15:
3773–83. h ps://doi.o g/10.1021/acs.jp o eome.6b00569 PMID: 27557128
49. Ba el J, Va ada ajan AR, Su a T, Ah ens CH, Maass S, Beche D. Op imized P o eomics Wo k low o
he De ec ion o Small P o eins. J P o eome Res. 2020; 19: 4004–18. h ps://doi.o g/10.1021/acs.
jp o eome.0c00286 PMID: 32812434
50. Swaney DL, Wenge CD, Coon JJ. Value o using mul iple p o eases o la ge-scale mass spec ome-
y-based p o eomics. J P o eome Res. 2010; 9: 1323–9. h ps://doi.o g/10.1021/p 900863u PMID:
20113005
51. Yuan P, D’Lima NG, Sla o SA. Compa a i e Memb ane P o eomics Re eals a Nonanno a ed E.coli
Hea Shock P o ein. Biochemis y. 2018; 57: 56–60. h ps://doi.o g/10.1021/acs.biochem.7b00864
PMID: 29039649
52. Yin X, Wu O M, Wang H, Hobbs EC, Shabalina SA, S o z G. The small p o ein Mg S and small RNA
Mg R modula e he Pi A phospha e sympo e o boos in acellula magnesium le els. Mol Mic obiol.
2019; 111: 131–44. h ps://doi.o g/10.1111/mmi.14143 PMID: 30276893
53. Fon aine F, Fuchs RT, S o z G. Memb ane localiza ion o small p o eins in Esche ichia coli. J Biol
Chem. 2011; 286: 32464–74. h ps://doi.o g/10.1074/jbc.M111.245696 PMID: 21778229
54. Zhou M, Boekho s J, F ancke C, Siezen RJ. Loca eP: genome-scale subcellula -loca ion p edic o o
bac e ial p o eins. BMC Bioin o ma ics. 2008; 9: 173. h ps://doi.o g/10.1186/1471-2105-9-173 PMID:
18371216
55. Meydan S, Ma ks J, Klepacki D, Sha ma V, Ba ano PV, Fi h AE, e al. Re apamulin-Assis ed Ribo-
some P o iling Re eals he Al e na i e Bac e ial P o eome. Mol Cell. 2019; 74: 481–93 e6. h ps://doi.
o g/10.1016/j.molcel.2019.02.017 PMID: 30904393
56. Impens F, Rolhion N, Radoshe ich L, Beca in C, Du al M, Mellin J, e al. N- e minomics iden i ies
P li42 as a memb ane minip o ein conse ed in Fi micu es and c i ical o s essosome ac i a ion in Lis-
e ia monocy ogenes. Na Mic obiol. 2017; 2: 17005. h ps://doi.o g/10.1038/nmic obiol.2017.5 PMID:
28191904
57. Ingolia NT, Ghaemmaghami S, Newman JR, Weissman JS. Genome-wide analysis in i o o ansla-
ion wi h nucleo ide esolu ion using ibosome p o iling. Science. 2009; 324: 218–23. h ps://doi.o g/10.
1126/science.1168978 PMID: 19213877
58. Go ochowski TE, Chelyshe a I, E iksen M, Nai P, Pede sen S, Igna o a Z. Absolu e quan i ica ion o
ansla ional egula ion and bu den using combined sequencing app oaches. Mol Sys Biol. 2019; 15:
e8719. h ps://doi.o g/10.15252/msb.20188719 PMID: 31053575
59. Ven u ini E, S ensson SL, Maaß S, Gelhausen R, Eggenho e F, Li L, e al. A global da a-d i en census
o Salmonella small p o eins and hei po en ial unc ions in bac e ial i ulence. mic oLi e. 2020;1.
60. Pe uschke H, Scho i C, Canzle S, Riesbeck S, Poehlein A, Daniel R, e al. Disco e y o no el commu-
ni y- ele an small p o eins in a simpli ied human in es inal mic obiome. Mic obiome. 2021; 9: 55.
h ps://doi.o g/10.1186/s40168-020-00981-z PMID: 33622394
61. D’Lima NG, Khi un A, Rosenbloom AD, Yuan P, Gassaway BM, Ba be KW, e al. Compa a i e P o eo-
mics Enables Iden i ica ion o Nonanno a ed Cold Shock P o eins in E.coli. J P o eome Res. 2017; 16:
3722–31. h ps://doi.o g/10.1021/acs.jp o eome.7b00419 PMID: 28861998
62. Bazzini AA, Johns one TG, Ch is iano R, Mackowiak SD, Obe maye B, Fleming ES, e al. Iden i ica ion
o small ORFs in e eb a es using ibosome oo p in ing and e olu iona y conse a ion. EMBO J.
2014; 33: 981–93. h ps://doi.o g/10.1002/embj.201488411 PMID: 24705786
63. Cassidy L, Helbig AO, Kaulich PT, Weidenbach K, Schmi z RA, Tholey A. Mul idimensional sepa a ion
schemes enhance he iden i ica ion and molecula cha ac e iza ion o low molecula weigh p o eomes
and sho open eading ame-encoded pep ides in op-down p o eomics. J P o eomics. 2021; 230:
103988. h ps://doi.o g/10.1016/j.jp o .2020.103988 PMID: 32949814
64. Omasi s U, Ah ens CH, Mulle S, Wollscheid B. P o e : in e ac i e p o ein ea u e isualiza ion and in e-
g a ion wi h expe imen al p o eomic da a. Bioin o ma ics. 2014; 30: 884–6. h ps://doi.o g/10.1093/
bioin o ma ics/b 607 PMID: 24162465
PLOS GENETICS
Small p o eins in S aphylococcus au eus
PLOS Gene ics | h ps://doi.o g/10.1371/jou nal.pgen.1009585 June 1, 2021 25 / 26