Full text
GigaScience, 7, 2018, 1–8
doi: 10.1093/gigascience/giy070
Ad ance Access Publica ion Da e: 12 June 2018
Da a No e
DATA NOTE
Me aMap: an a las o me a ansc ip omic eads in
human disease- ela ed RNA-seq da a
L.M. Simon 1,*,S.Ka g
1, A.J. Wes e mann2,3,M.Engel
1,4, A.H.A. Elbehe y5,
B. Hense1,M.Heinig
1,L.Deng
5and F.J. Theis 1,6,*
1Helmhol z Zen um M ¨
unchen, Ge man Resea ch Cen e o En i onmen al Heal h, Ins i u e o
Compu a ional Biology, Neuhe be g, Ge many, 2Ins i u e o Molecula In ec ion Biology, Uni e si y o
W¨
u zbu g, W ¨
u zbu g, Ge many, 3Helmhol z Ins i u e o RNA-Based In ec ion Resea ch, W¨
u zbu g, Ge many,
4Helmhol z Zen um M ¨
unchen, Ge man Resea ch Cen e o En i onmen al Heal h, Scien i ic Compu ing
Resea ch Uni , Neuhe be g, Ge many, 5Helmhol z Zen um M ¨
unchen, Ge man Resea ch Cen e o
En i onmen al Heal h, Ins i u e o Vi ology, Neuhe be g, Ge many and 6Depa men o Ma hema ics,
Technische Uni e si ¨
a M ¨
unchen, Munich, Ge many
∗Co espondence add ess. L.M. Simon, Helmhol z Zen um M ¨
unchen Ge man Resea ch Cen e o En i onmen al Heal h, Ins i u e o Compu a ional
Biology, Ingols ¨
ad e Lands aße, 185764, Neuhe be g; E-mail: [email p o ec ed] h p://o cid.o g/0000-0001-6148-8861; F.J. Theis;
E-mail: abian. heis@helmhol z-muenchen.de h p://o cid.o g/0000-0002-2419-1943
Abs ac
Backg ound: Wi h he ad en o he age o big da a in bioin o ma ics, la ge olumes o da a and high-pe o mance
compu ing powe enable esea che s o pe o m e-analyses o publicly a ailable da ase s a an unp eceden ed scale. E e
mo e s udies imply he mic obiome in bo h no mal human physiology and a wide ange o diseases. RNA sequencing
echnology (RNA-seq) is commonly used o in e global euka yo ic gene exp ession pa e ns unde de ined condi ions,
including human disease- ela ed con ex s; howe e , i s gene ic na u e also enables he de ec ion o mic obial and i al
ansc ip s. Findings: We de eloped a bioin o ma ic pipeline o sc een exis ing human RNA-seq da ase s o he p esence o
mic obial and i al eads by e-inspec ing he non-human-mapping ead ac ion. We alida ed his app oach by
ecapi ula ing ou comes om six independen , con olled in ec ion expe imen s o cell line models and compa ed hem
wi h an al e na i e me a ansc ip omic mapping s a egy. We hen applied he pipeline o close o 150 e aby es o publicly
a ailable aw RNA-seq da a om mo e han 17,000 samples om mo e han 400 s udies ele an o human disease using
s a e-o - he-a high-pe o mance compu ing sys ems. The esul ing da a om his la ge-scale e-analysis a e made
a ailable in he p esen ed Me aMap esou ce. Conclusions: Ou esul s demons a e ha common human RNA-seq da a,
including hose a chi ed in public eposi o ies, migh con ain aluable in o ma ion o co ela e mic obial and i al
de ec ion pa e ns wi h di e se diseases. The p esen ed Me aMap da abase hus p o ides a ich esou ce o hypo hesis
gene a ion owa d he ole o he mic obiome in human disease. Addi ionally, codes o p ocess new da ase s and pe o m
s a is ical analyses a e made a ailable.
Keywo ds: high-pe o mance compu ing; big da a; RNA-seq; sequence ead a chi e; me a ansc ip omics; mic obiome;
i ome; human disease; in ec ion
Recei ed: 22 Feb ua y 2018; Re ised: 1 June 2018; Accep ed: 4 June 2018
C
The Au ho (s) 2018. Published by Ox o d Uni e si y P ess. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons
A ibu ion License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium,
p o ided he o iginal wo k is p ope ly ci ed.
1
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
2Me aMap
Da a Desc ip ion
Con ex
Recen s udies ha e demons a ed he pa amoun impo ance
o he mic obiome o human heal h and disease [1]. Fo exam-
ple, imbalance o he human gu mic obiome was linked o non-
communicable diseases such as obesi y [2,3], diabe es [4], ca -
dio ascula disease [5], ch onic obs uc i e pulmona y disease
[6], and colo ec al ca cinoma [7,8], o name jus a ew.
The ad en o high- h oughpu sequencing echnologies has
e olu ionized he li e sciences. RNA sequencing (RNA-seq)
echnology p oduces one o he mos equen nex -gene a ion
sequencing da a ypes and has been applied o he s udy o a
la ge numbe o biological samples ele an o human disease.
The majo i y o he unde lying aw da a a e eely accessible
om da a eposi o ies such as he Gene Exp ession Omnibus
(>1,700 human RNA-seq da ase s as o Janua y 2018) and he Se-
quence Read A chi e (SRA) [9].
Howe e , hese da a a e ypically exclusi ely used o single
species (i.e., human) ansc ip omics such as di e en ial gene
exp ession and al e na i e splicing analysis [9,10]. Reads ha
do no map on o he human genome a e conside ed noise o
con amina ion and he e o e a e gene ally igno ed [11,12](col-
lec i ely abou 9% o o al eads, Fig. 1). Fi e yea s ago, i was pos-
ula ed ha in e species in e ac ions migh be s udied by simul-
aneous de ec ion and quan i ica ion o RNA ansc ip s om a
gi en hos and a mic obe ia “dual” RNA-seq [13]. Meanwhile,
his app oach has been success ully applied o he in e ac ion o
mammalian cells wi h di e se bac e ial [14] and i al pa hogens
[15-19].
Inspi ed by dual RNA-seq, in his s udy we hypo hesize ha
eads in a chi ed RNA-seq da ase s de i ed om human p i-
ma y cells o issue samples ha ail o map agains he hu-
man e e ence genome may con ain aluable in o ma ion abou
hep esenceo ce ainmic obesin he espec i ebodyniches
and/o unde de ined disease condi ions. To enable me a an-
sc ip omic s udy o hese da a, we combined exis ing ead align-
men and me agenomic classi ica ion so wa e in o a wo-s ep
“omni” RNA-seq pipeline o comp ehensi ely quan i y a chaeal,
bac e ial, and i al eads in human RNA-seq da a (Fig.1).
In he i s s ep o his so-called Me aMap pipeline, all eads
a e aligned agains he human genome using he ul a- as RNA-
seq aligne Spliced T ansc ip s Alignmen o a Re e ence so -
wa e (STAR) [20]. Subsequen ly, only he ac ion o unmapped
eads is subjec ed o me a ansc ip omic classi ica ion using
CLARK-S [21] (see Me hods o de ails). The combina ion be-
ween scalabili y and accu acy was he main mo i a ion behind
choosing hese wo so wa e packages o e compe ing me h-
ods [22,23]. I is impo an o no e ha CLARK-S uses a se
o uniquely disc imina i e sho sequences a he species le el
o classi y eads. The e o e, eads con aining nondisc imina i e
sequences ha ail o be uniquely assigned o a single species,
e.g., eads o igina ing om he bac e ial ibosomal 16S RNA
gene will be conside ed “unclassi ied” (al oge he 8.6% in Fig.1).
The ou pu o CLARK-S is an ope a ional axonomic uni s
(OTUs) coun ma ix, whe e ows co espond o i al, bac e ial,
and a cheal species and columns co espond o (human) sam-
ples. Each en y co esponds o he numbe o non-human eads
classi ied o he espec i e species. Fo con enience, in he ol-
lowing, we e e o he se o mic obial and i al species p o iled
using ou app oach as “me a ea u es.”
By sc eening he s udy abs ac s o he SRA o sea ch
e ms p io i izing human clinical da ase s de i ed om polyA-
independen sequencing p o ocols (see Me hods), we iden i-
ied mo e han 400 s udies ele an o human disease comp is-
ing mo e han 17,000 cDNA lib a ies (close o 150 e aby es o
aw sequencing da a). Raw sequencing eads om hese s ud-
ies we e downloaded and analyzed using he high-pe o mance
compu ing sys em o he Leibniz Supe compu ing Cen e (LRZ)
o he Ba a ian Academy o Sciences and Humani ies, which a-
cili a ed ul a- as p ocessing wi h median speeds o 25 and 21
million eads pe hou pe co e pe un o he STAR and CLARK-
S s eps, espec i ely. O he mo e han 500 billion RNA-seq eads
p ocessed, a ound 91% could be mapped o he human genome.
A ac ion o 8.6% o all eads emained nondisc imina i e a he
species le el and de ined as “unclassi ied.” In addi ion, 0.03%,
0.20%, and 0.39% o all eads we e assigned o a chaeal, bac e-
ial, o i al me a ea u es, espec i ely. Despi e hese ela i ely
low pe cen ages, he absolu e numbe s o eads classi ied we e
in he hund ed millions o billions, enabling s a is ical analyses.
Me hods
High-pe o mance compu ing en i onmen
P ojec compu a ions including download, alignmen o eads
on o he human genome, and me a ea u e quan i ica ion we e
made on he high-pe o mance Linux Clus e a he LRZ [24].
RNA-seq da a e ie al
Raw nex -gene a ion sequencing da a we e downloaded om
he SRA. The R package SRAdb was downloaded on 23 May 2017
and used o que y he SRA da abase. To iden i y SRA p ojec s
ha con ain ansc ip omic analyses o human RNA-seq da a,
he SRA a ibu es “ axon id,” “lib a y sou ce,” “lib a y s a egy,”
and “pla o m” we e sea ched o he e ms “9606,” “TRAN-
SCRIPT,” “RNA-seq,” and “ILLUMINA,” espec i ely. To emo e
po en ial bias de i ed om di e en sequencing echnologies,
we also es ic ed he que y o SRA uns anno a ed wi h “ILLU-
MINA” in SRA a ibu e “pla o m.” To exclude s udies wi h in-
su icien sample size o s a is ical analysis, he que y was e-
s ic ed o SRA p ojec s con aining mo e han i e uns. To a oid
concen a ing he analysis on a small numbe o la ge p ojec s,
he que y was es ic ed o SRA p ojec s wi h ewe han 500
uns. To iden i y s udies ocusing on pheno ypes ele an o hu-
man disease, we es ic ed he que y o uns con aining a leas
one o mo e o he e ms “disease,” “pa ien ,” “p ima y,” and
“clinical” in he SRA a ibu e “s udy abs ac .” To exclude in i o
(cell-cul u e) expe imen s bu ocus on p ima y (clinical) sam-
ples, SRA uns con aining he e ms “mu an ” o “cell-line” we e
emo ed om ou selec ion. Fu he mo e, SRA uns con aining
he e ms “single cell” and “GTEx” we e emo ed. Finally, sam-
ples wi h ewe han 1 million o al eads o ead leng hs <50 bp
we e excluded. The desc ibed que y esul ed in 484 sho ead
p ojec s (SRPs) con aining 21,659 RNA-seq uns. Due o echnical
p oblems (i.e., missing URLs, es ic ed access), we we e unable
o download a ac ion o 4,078 samples.
Human alignmen
Alignmen o eads agains he human e e ence genome (hg38)
and simul aneous human gene exp ession quan i ica ion was
conduc ed wi h STAR ( e sion 2.5.2). To inc ease mapping speed
o a la ge numbe o samples, we used he –genomeLoad LoadAnd-
Keep unc ion o load he STAR index once and keep i in
memo y o subsequen alignmen s. The pa ame e –quan mode
GeneCoun s was used o gene a e he human gene exp es-
sion coun ables. Unmapped eads we e sa ed wi h he –
ou ReadsUnmapped Fas x pa ame e . To u he inc ease mapping
speed, mul iple h eads we e used as implemen ed wi h he pa-
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
Simon e al. 3
Figu e 1: Schema ic o he Me aMap pipeline. Mo e han 400 p ojec s om s udies ele an o human disease we e iden i ied in he SRA da abase. Mo e han 500 billion
RNA-seq eads we e downloaded and i s il e ed by mapping hem on o he human genome. The emaining eads unde wen me a ea u e classi ica ion. I is no ed
ha 90.7% o all eads mapped o he human genome; 0.03%, 0.20%, and 0.39% o all eads we e assigned o a chaeal, bac e ial, o i al me a ea u es, espec i ely; and
8.6% o all eads emain nondisc imina i e a he species le el (“unclassi ied”).
ame e – unTh eadN 28. Runs wi h ewe han 30% eads map-
ping o he human genome we e excluded om downs eam
analysis. All human alignmen s we e conduc ed on he LRZ
“CoolMUC2” Linux-Clus e . This clus e con ains 384 nodes wi h
64 GB andom access memo y (RAM) memo y and 28 co es each.
Me a ea u e quan i ica ion
Me a ea u e quan i ica ion was conduc ed wi h CLARK-S ( e -
sion 1.2.3). CLARK-S is a so wa e me hod o as and accu-
a e sequence classi ica ion o me agenomic nex -gene a ion
sequencing da a, including RNA-seq da a. One majo issue du -
ing he classi ica ion o me agenomic da a is he ising numbe
o a ge s o align agains . CLARK-S sol es his issue by build-
ing a la ge index ile consis ing o disc imina i e k-me s. The
me agenomic e e ence da abase was gene a ed ollowing he
desc ip ion o he CLARK websi e using he ollowing wo com-
mands: se a ge s.sh bac e ia i us –species and buildSpacedDB.sh.
This da abase con ained 16,551 genome sequences co espond-
ing o 6,979 unique species (Addi ional File 2). To allow uni-
o m p ocessing, pai ed-end sequencing expe imen s we e an-
alyzed independen ly. Each single unmapped ead ile was
used as inpu o CLARK-S wi h he ollowing pa ame e s: clas-
si y me agenome.sh –spaced –O lis o FASTQ iles. To inc ease
classi ica ion speed, he CLARK-S exp ess mode was selec ed
and mul iple h eads we e used wi h pa ame e s –m 2 and –n
32, espec i ely. The ou pu iles o his s ep con ain all inpu
ead iden i ie s wi h he co esponding me a ea u e classi ica-
ion. In he subsequen s ep, o al coun s a e summa ized o
each ea u e wi h he es ima e abundance.sh command. To en-
able compa ison ac oss single-end and pai ed-end expe imen s,
me a ea u e coun s om pai ed-end expe imen s we e a e -
aged and subsequen ly ounded o conse e coun dis ibu ion.
To accoun o a ying sequencing dep hs, me a ea u e abun-
dance was es ima ed as he numbe o eads pe million o-
al eads sequenced. Me a ea u e quan i ica ion was conduc ed
on he LRZ “Te amem” Linux-Clus e . This clus e con ains one
node wi h 6,144 GB RAM memo y and 96 co es.
BLAST-based me a ea u e classi ica ion
To alida e esul s gene a ed by he Me aMap pipeline, he Basic
Local Alignmen Sea ch Tool (BLAST) [25] was used as ollows. A
BLAST da abase was c ea ed om he same genome sequences
used in he CLARK-S app oach. Then, eads we e aligned o
his da abase using BLASTN wi h a h eshold E- alue o 1e-10.
P oduced coun s om pai ed-end expe imen s we e a e aged.
Fo each ile, BLAST was done by unning app oxima ely 10 kb
chunks ( eco d sepa a o “>”) in pa allel using pa allel [41] (28
jobs), each wi h eigh h eads using one node on he LRZ “Cool-
MUC3” Linux Clus e . This clus e con ains 148 nodes wi h 96
GB RAM memo y and 64 co es each. Ou pu was pa sed o ex-
clusi ely keep eads ha could be assigned a he species le el.
Di e en ial me a ea u e abundance
Di e en ial me a ea u e abundance analysis was pe o med us-
ing he R package DESeq2 [26]. DESeq2 models di e en ial gene
exp ession by i ing a nega i e binomial dis ibu ion o he aw
coun s unde lying RNA-seq da a. This amewo k can accoun
o con ounding a iables such as sequencing dep h. The e o e,
he da a need no be no malized p io o s a is ical in e -sample
compa isons. Fo each o he ou published bona ide dual RNA-
seq s udies, we classi ied samples in o he ollowing wo g oups
based on he p o ided anno a ions: samples expec ed o con ain
he known pa hogen, such as human papilloma i us-posi i e
umo s in he Zhang e al. s udy [28], and pa hogen- ee con-
ols, such as mock- ea ed cells in he Wes e mann e al. [27]
s udy. Using his bina y ou come, we pe o med di e en ial ex-
p ession analysis ac oss all de ec ed me a ea u es. To accoun
o sequencing dep h, lib a y size ac o s we e es ima ed om
he o al numbe o sequenced eads. The dispe sion o he neg-
a i e binomial dis ibu ion was es ima ed using a local linea e-
g ession as implemen ed in he DESeq() unc ion ia he i Type
pa ame e “local.”
Da a alida ion and quali y con ol
We alida ed ou app oach by eco e ing he g ound u h in
bona ide dual RNA-seq expe imen s pe o med wi h human
cell lines and samples om pa ien s wi h well-known in ec-
ion s a us. O he ou selec ed s udies, one analyzed an in ec-
ion model based on a bac e ial (Salmonella en e ica se o a Ty-
phimu ium) and h ee based on dis inc i al pa hogens (human
papilloma i us, he pes simplex i us, hino i us). As expec ed,
Me aMap de ec ed he known pa hogen a highe le els in he
espec i e s udy compa ed o he o he s udies and pa hogens
(Table1). Howe e , compa isons ac oss s udies and me a ea u es
may be biased by echnical con ounde s (discussed in de ail in
he Re-use po en ial sec ion). The e o e, we ocused ou analy-
sis on he compa ison o a single me a ea u e ac oss subjec s
wi hin a s udy. Using he anno a ion p o ided in he espec i e
s udy, we pe o med di e en ial me a ea u e abundance analy-
sis o iden i y hose me a ea u es ha show he la ges ela i e
di e ence in abundance le els be ween he in ec ed and con ol
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
4Me aMap
Table 1: O e iew o ou dual RNA-seq s udies used o alida e he Me aMap pipeline.
S udy
In ec ion
agen
To al
eads
Salmonella
en e ica
Alphapapilloma i us
9
Human
alphahe pes i us
1
Rhino i us
A
Wes e mann
e al. [27]
Salmonella en e ica se o a
Typhimu ium
1.0e+07 6.3e+03 1.2e-01 1.5e-01 1.2e-01
Zhang e al.
[28]
Human papilloma i us 4.6e+07 3.0e-02 5.1e+01 2.2e-02 2.2e-02
Ru kowski e
al. [29]
He pes simplex i us 3.5e+07 1.1e+00 3.1e-02 3.1e+04 3.0e-02
Bai e al. [30] Rhino i us 6.6e+06 2.0e-01 1.5e-01 1.5e-01 4.4e+01
To al eads column depic s he a e age ead dep h pe sample o each s udy. A e age me a ea u e abundance o alphapapilloma i us 9, Salmonella en e ica, human
alphahe pes i us 1, and hino i us A a e shown in eads pe million. The co ec in ec ion agen o he espec i e s udy is highligh ed in bold on
0
25
50
75
0510
−log10 p− alue
Wes e mann e al
0
5000
10000
15000
in ec ed
(n=36)
mock
(n=6)
Salmonella en e ica
A
0
10
20
04812
Zhang e al
Alphapapilloma i us 9
B
0.0
5.0
10.0
−5.0 −2.5 0.0 2.5 5.0 7.5
Ru kowski e al
0
25000
50000
75000
in ec ed
(n=8)
mock
(n=2)
H. Alphahe pes i us 1
C
0
20
40
048
Fold change (log2)
Bai e al
0
50
100
150
200
in ec ed
(n=12)
ehicle
(n=12)
Rhino i us A
D
0
100
200
300
400
500
nega i e
(n=18)
posi i e
(n=18)
Figu e 2: Di e en ial me a ea u e abundance analysis o con olled in ec ion expe imen s eco e s g ound u h. “Volcano” plo s show old change and in e ed P alue
on he xand yaxes, espec i ely. Each do ep esen s a me a ea u e. The mos signi ican me a ea u e is colo ed in ed. Inse s display box plo s o he abundance
le els in eads pe million o he op hi me a ea u e ac oss condi ions o each s udy. Fo all box plo s, he box ep esen s he in e qua ile ange, he ho izon al line
in he box is he median, and he whiske s ep esen 1.5 imes he in e qua ile ange.
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
Simon e al. 5
Figu e 3: Analysis o lymphoblas cell line expe imen s u he suppo s he Me aMap pipeline. (A and B) Mean abundance le els ac oss all samples o he op i e
me a ea u es o p ojec s SRP041338 and SRP091453, espec i ely. (C) Rela i e p opo ion o eads mapping o EBV, phiX, and all o he me a ea u es ac oss RNA-seq
samples. (D) Cumula i e dis ibu ion plo o he a e age p opo ion o bac e ial me a ea u e eads ac oss all p ojec s. Pu ple and pink e ical lines highligh p ojec s
SRP041338 and SRP091453, espec i ely.
samples (see Me hods o de ails). The co ec in ec ion agen
showed he mos signi ican di e ence ac oss all me a ea u es
be ween in ec ed and con ol samples o each s udy (Fig.2). Fo
example, Wes e mann e al. [27] gene a ed dual RNA-seq da a
om HeLa cells in ec ed wi h he en e ic bac e ial pa hogen
S. en e ica se o a Typhimu ium and compa ed hem o mock-
ea ed con ol samples. Acco dingly, we obse ed S. en e ica as
he mos di e en ially abundan me a ea u e be ween he in-
ec ed and he con ol samples (P<1e-75, Fig. 2A). Likewise, we
eco e ed alphapapilloma i us 9,human alphahe pes i us 1 (also
known as he pes simplex i us 1), and hino i us A as he mos
di e en ially abundan me a ea u es in he da a om Zhang e
al. [28], Ru kowski e al. [29], and Bai e al. [30], espec i ely. In he
Wes e mann e al. [27] and Ru kowski e al. [29] s udies, se e al
addi ional me a ea u es showed a s ong di e en ial abundance
e ec (Fig. 2A and 2C). These me a ea u es we e closely ela ed
o he ue in ec ion agen , i.e., Salmonella bongo i (P<1e-67) and
Panine alphahe pes i us 3 (P<1e-9) o he Wes e mann e al. [27]
o Ru kowski e al. [29] s udy, espec i ely. These indings con-
i m ha ou Me aMap pipeline ecapi ula es esul s om dedi-
ca ed dual RNA-seq s udies, i.e., s udies based on known in ec-
ious agen s. The e o e, Me aMap may be equally sui ed o de-
ec p e iously unknown mic obial and i al species in human
p ima y samples.
As an addi ional con ol, we e-analyzed wo p ojec s con-
ained in ou da a collec ion ha a e de i ed om he B lym-
phoblas cell line unde nonin ec ious condi ions. Howe e ,
since Eps ein-Ba i us (EBV) is used o ans ec ion and ans-
o ma ion o lymphocy es o lymphoblas s, we expec ed o de-
ec eads om his i us in hese p ojec s [31], bu no u -
he i al o mic obial eads [32]. Indeed, he mos abundan
me a ea u es in each p ojec we e domina ed by eads classi-
ied o gammahe pes i us 4 (also known as EBV) and En e obac e-
ia phage phiX174 sensu la o (phiX), commonly used as spike-in
in Illumina sequencing uns [33](Fig.3A and 3B). On a e age,
95% and 97% o all me a ea u e eads we e classi ied as phiX o
EBV o p ojec s SRP041338 and SRP091453, espec i ely (Fig. 3C).
Con e sely, he abundance o eads mapping o bac e ial species
o hese wo p ojec s co esponds o he bo om pe cen ile as
compa ed o all o he p ojec s in he Me aMap da abase, sup-
po ing s e ili y o his cell line (Fig. 3D). This demons a es ha
Me aMap no only is capable o edisco e ing known pa hogenic
species ( ue posi i es) in con olled in ec ion expe imen s (Fig.
2) bu i also minimizes he de ec ion o alse posi i es o , a
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
6Me aMap
0 5000 10000 15000
0 5000 10000 15000
Salmonella en e ica le els
CLARK-S based abundance (RPM)
BLAST based abundance (RPM)
Spea man
Rho 0.98
P < 2.2e-16
1 5 50 500 5000
1 10 100 1000 10000
A e age me a ea u e abundance
CLARK-S based (RPM)
BLAST based (RPM)
Spea man
Rho 0.16
P 3.1e−10
AB
C
Salmonella en e ica
CLARK−S
BLAST
0246
Reads pe hou pe h ead
(log10)
8
Figu e 4: Al e na i e BLAST-based classi ica ion me hod alida es me a ea u e abundance es ima es by Me aMap. (A) A e age me a ea u e eads pe million le els
de i ed using he CLARK-S so wa e, as implemen ed in he Me aMap pipeline, and a BLAST-based al e na i e app oach on he xand yaxes, espec i ely. (B) Co ela ion
in S. en e ica abundance le els be ween he wo classi ica ion app oaches. (C) Di e ence in classi ica ion speed be ween he BLAST and CLARK-S me a ansc ip omic
classi ica ion. The yaxis shows he numbe o eads p ocessed pe hou pe h ead in log10 space.
leas , p o ides measu es such as abundance and signi icance,
allowing he use o iden i y and coun e selec hose species.
As a echnical alida ion, we compa ed ou app oach o an
al e na i e me a ansc ip omic classi ica ion s a egy o he
Wes e mann e al. [34] s udy. All non-human eads we e aligned
using BLASTN o a BLAST da abase consis ing o he same ge-
nomic sequences used by CLARK-S (see Me hods o de ails). The
a e age me a ea u e abundances ac oss all 42 samples de i ed
om he BLAST-based app oach and CLARK-S co ela ed sig-
ni ican ly (Spea man co ela ion, Rho: 0.16, P: 3.1e-10) (Fig. 4A).
BLAST showed highe sensi i i y and de ec ed mo e me a ea-
u es compa ed o CLARK-S (indica ed by he accumula ion o
do s a alue 0 on he xaxis in Fig. 4A).Thisismos lyobse ed
o low abundance me a ea u es ha could ep esen low coun s
de i ed om sequencing and/o mapping e o s. Howe e , mos
impo an ly, he ue pa hogen me a ea u e “Salmonella en e -
ica” showed e y high co ela ion ac oss samples be ween he
BLAST- and CLARK-based abundance es ima es (Fig. 4B). No e-
wo hy, he Me aMap pipeline p ocessed eads mo e han h ee
o de s o magni ude as e han BLAST, demons a ing a sig-
ni ican speed ad an age while gene a ing compa able esul s
(Fig. 4C).
Re-use po en ial
Mic obial and i al con amina ion in nex -gene a ion sequenc-
ing da a has been obse ed. I can be caused by mapping e o s
due o genome sequence simila i y be ween di e en species
[35,36]. In addi ion, echnical con ounde s can obs uc he
analysis and po en ially gene a e a i icial di e ences i no con-
side ed p ope ly. Fo example, di e en ypes o human sam-
ples may con ain di e en amoun s o non-human ma e ial
due o a ying s e ili y o he issues. Fu he mo e, sequencing
dep h may in oduce a de ec ion loo o me a ea u es ha a e
no abundan . The e o e, compa isons ac oss di e en issues
and sequencing dep hs may gene a e a i icial di e ences. Addi-
ionally, gi en ha only uniquely disc imina i e sequences a e
coun ed, he absolu e abundance le els may no be compa a-
ble ac oss me a ea u es. Finally, he Me aMap pipeline cap u es
me a ea u e abundance a he RNA le el, which may no nec-
essa ily co espond o genomic abundance le els. Me a ea u es
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
Simon e al. 7
may no be abundan a he DNA le el bu highly ansc ip ion-
ally ac i e and hus abundan ly de ec ed a he RNA le el, o he
in e se. These po en ial challenges need o be aken in o con-
side a ion when compa ing ac oss me a ea u es.
To minimize hese e ec s, we encou age ocusing on s udies
ha include in ap ojec compa isons ha es one me a ea u e
a a ime, as exempli ied in he di e en ial me a ea u e abun-
dance analysis. Ou a ionale is ha echnical con ounde s, in
con as o biologically meaning ul changes, should a ec all
uns wi hin a p ojec o he same ex en and he e o e no show
condi ion-speci ic e ec s. Fo example, in he Wes e mann e al.
s udy [34], we de ec ed subs an ial le els o phiX in bo h condi-
ions (in ec ed samples and mock- ea ed con ols), bu only he
“Salmonella” me a ea u e showed a condi ion-speci ic e ec . We
aim o add ess he challenges inhe en o in e p ojec and in e -
me a ea u e compa isons in u u e wo k.
All he aw da a desc ibed in he p esen s udy we e pub-
licly a ailable, ye ha e been e y cumbe some o ex ac in-
di idually. The p esen ed Me aMap da abase makes hese da a
easily accessible o a e y b oad communi y, he eby allow-
ing o global compa isons o e hund eds o indi idual s udies
and housands o sampled condi ions. While we a emp ed o
minimize he isk o de ec ing alse posi i es (Fig. 3), i should
be no ed ha no all me a ea u es classi ied by Me aMap will
necessa ily e e o ue biological ac o s. No ewo hy, ou ap-
p oach e eals a co ela ion be ween me a ea u es and disease,
no causali y, and canno disc imina e disease-associa ed e -
ec s om po en ial ea men e ec s. Howe e , ou pipeline
p o ides he use wi h a scien i ic s a ing g ound o alida e he
p esence/absence o de ined mic obial and i al species unde
de ined condi ions and explo e he unde lying biology and sig-
ni icance in g ea e de ail. As a po en ial use case o hese da a,
use s can es o associa ions o mic obial o i al me a ea-
u es wi h a ple ho a o human diseases o be ween hemsel es.
In addi ion, use s wi h in e es in a speci ic bac e ial o i al
species can easily iden i y s udies and, consequen ly, disease
con ex s in which eads om his o ganism we e de ec ed. This
could gi e an impo an i s hin o assess whe he he espec-
i e species migh be implica ed in a gi en human disease e i-
ology. Fu he mo e, his esou ce p o ides he oppo uni y o
suppo indings de i ed om s anda d mic obiome p o iling
echnologies, such as 16S RNA gene based o sho gun me age-
nomics [37]. Finally, me a ea u e de ec ion in human clinical
RNA-seq samples may p o ide a diagnos ic ad an age when
s udying mic obes o i uses ha a e challenging o isola e.
The composi e me a ea u e OTU coun able, de i ed om
17 278 cDNA lib a ies om 436 SRA p ojec s, including anno a-
ions is p o ided o download [38].
A ailabili y o sou ce code and equi emen s
P ojec name: Me aMap
P ojec home: h ps://gi hub.com/ heislab/Me aMap
Ope a ing sys em(s): Pla o m-independen
P og amming language: Unix command line, R
O he equi emen s: STAR and CLARK-S may equi e la ge
amoun s o memo y (>100 GB)
License: GNU GPL
A ailabili y o suppo ing da a
The da ase s suppo ing he esul s p esen ed he e a e a ailable
in he GigaScience Da abase eposi o y [38]. The p o ocols a e also
a ailable a [40].
Addi ional ile
Addi ional File1.cs
Abb e ia ions
BLAST: basic local alignmen sea ch ool; EBV: Eps ein-Ba i us;
LRZ: Leibniz Supe compu ing Cen e; OTU: ope a ional axo-
nomic uni ; phiX: En e obac e ia phage phiX174 sensu la o; RNA-
seq: RNA sequencing; SRA: Sequencing Read A chi e; SRP: sho
ead p ojec ; STAR: Spliced T ansc ip s Alignmen o a Re e ence
so wa e.
Compe ing in e es s
The au ho s decla e ha hey ha e no compe ing in e es s.
Funding
L.S. acknowledges unding om he Eu opean Union’s Ho i-
zon 2020 Resea ch and Inno a ion P og amme unde he Ma ie
Sklodowska-Cu ie g an ag eemen (753039). The ope a ion o
he LRZ Linux Clus e is unded ia he Ba a ian S a e Minis y
o Educa ion, Science, and he A s.
Au ho con ibu ions
Concep ualiza ion: L.S., M.E., L.D., and B.H.; o mal analysis: L.S.,
M.H., S.K., and A.E.; in es iga ion: L.S., A.J.W., and M.E.; me hod-
ology: L.S., S.K, M.H.; w i ing he o iginal d a : L.S. and A.J.W;
w i ing, e iewing, and edi ing: L.S., A.J.W., M.E., A.E., L.D., M.H.,
and F.T.; supe ision: L.D., M.H., and F.T.
Acknowledgmen s
The au ho s hank Yu Wang and Fe dinand Jami zky om he
LRZ o hei suppo .
Re e ences
1. Young VB. The ole o he mic obiome in human heal h and
disease: an in oduc ion o clinicians. BMJ 2017;356:j831.
2. Tu nbaugh PJ, Ley RE, Mahowald MA, e al. An obesi y-
associa ed gu mic obiome wi h inc eased capaci y o en-
e gy ha es . Na u e 2006;444:1027–31.
3. Henao-Mejia J, Elina E, Jin C, e al. In lammasome-media ed
dysbiosis egula es p og ession o NAFLD and obesi y. Na u e
2012;482:179–85.
4. Cani PD, Bibiloni R, Knau C, e al. Changes in gu mic o-
bio a con ol me abolic endo oxemia-induced in lamma ion
in high- a die -induced obesi y and diabe es in mice. Dia-
be es 2008;57:1470–81.
5. Wang Z, Klip ell E, Benne BJ, e al. Gu lo a me abolism
o phospha idylcholine p omo es ca dio ascula disease. Na-
u e 2011;472:57–63.
6. Engel M, Endes elde D, Schlo e -Hai B, e al. In luence o lung
CT changes in ch onic obs uc i e pulmona y disease (COPD)
on he human lung mic obiome. PLoS One 2017;12:e0180859.
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
8Me aMap
7. Kos ic AD, Ge e s D, Pedamallu CS, e al. Genomic analysis
iden i ies associa ion o Fusobac e ium wi h colo ec al ca ci-
noma. Genome Res 2012;22:292–8.
8. Cas ella in M, Wa en RL, F eeman JD, e al. Fusobac e ium nu-
clea um in ec ion is p e alen in human colo ec al ca cinoma.
Genome Res 2012;22:299–306.
9. Kodama Y, Shumway M, Leinonen R, In e na ional Nu-
cleo ide Sequence Da abase Collabo a ion. The Sequence
Read A chi e: explosi e g ow h o sequencing da a. Nucleic
Acids Res 2012;40:D54–6.
10.Conesa A, Mad igal P, Ta azona S, e al. A su ey o bes p ac-
ices o RNA-seq da a analysis. Genome Biol 2016;17:13.
11.Gouin A, Legeai F, Nouhaud P, e al. Whole-genome e-
sequencing o non-model o ganisms: lessons om un-
mapped eads. He edi y 2015;114:494–501.
12.Peng X, Wang J, Zhang Z, e al. Re-alignmen o he un-
mapped eads wi h base quali y sco e. BMC Bioin o ma ics
2015;16(Suppl 5):S8.
13.Wes e mann AJ, Go ski SA, Vogel J. Dual RNA-seq o pa hogen
and hos . Na Re Mic obiol 2012;10:618–30.
14.Wes e mann AJ, Ba quis L, Vogel J. Resol ing hos -pa hogen
in e ac ions by dual RNA-seq. PLoS Pa hog 2017;13:e1006033.
15.Ju anic Lisnic V, Babic Cac M, Lisnic B, e al. Dual analysis o
he mu ine cy omegalo i us and hos cell ansc ip omes e-
eal new aspec s o he i us-hos cell in e ace. PLoS Pa hog
2013;9:e1003611.
16.Xu G, S ong MJ, Lacey MR, e al. RNA CoMPASS: a dual ap-
p oach o pa hogen and hos ansc ip ome analysis o RNA-
seq da ase s. PLoS One 2014;9:e89445.
17.Pa k S-J, Kuma M, Kwon H-I, e al. Dynamic changes in hos
gene exp ession associa ed wi h H5N8 a ian in luenza i us
in ec ion in mice. Sci Rep 2015;5:16512.
18.Saxena K, Simon LM, Zeng X-L, e al. A pa adox o ansc ip-
ional and unc ional inna e in e e on esponses o human
in es inal en e oids o en e ic i us in ec ion. P oc Na l Acad
Sci 2017;114:E570–9.
19.Wesolowska-Ande sen A, E e man JL, Da idson R, e al. Dual
RNA-seq e eals i al in ec ions in as hma ic child en wi h-
ou espi a o y illness which a e associa ed wi h changes in
he ai way ansc ip ome. Genome Biol 2017;18:12.
20.Dobin A, Da is CA, Schlesinge F, e al. STAR: ul a as uni e -
sal RNA-seq aligne . Bioin o ma ics 2012;29:15–21.
21.Ouni R, Lona di S. Highe classi ica ion sensi i i y o
sho me agenomic eads wi h CLARK-S. Bioin o ma ics
2016;32:3823–5.
22.Lindg een S, Adai KL, Ga dne PP. An e alua ion o he ac-
cu acy and speed o me agenome analysis ools. Sci Rep
2016;6:19233.
23.Engs ¨
om PG, S eijge T, Sipos B, e al. Sys ema ic e alua ion
o spliced alignmen p og ams o RNA-seq da a. Na Me h-
ods 2013;10:1185–91.
24. www.l z.de/se ices/compu e/linux-clus e , Leibniz Supe -
compu ing Cen e
25.Al schul S. Basic local alignmen sea ch ool. J Mol Biol
1990;215:403–10.
26.Lo e MI, Hube W, Ande s S. Mode a ed es ima ion o old
change and dispe sion o RNA-seq da a wi h DESeq2 [In e -
ne ]. Genome Biol 2014;15:550.
27.Wes e mann AJ, F¨
o s ne KU, Amman F, e al. Dual RNA-seq
un eils noncoding RNA unc ions in hos –pa hogen in e ac-
ions. Na u e 2016;529:496–501.
28.Zhang Y, Kone a LA, Vi ani S, e al. Sub ypes o HPV-posi i e
head and neck cance s a e associa ed wi h HPV cha ac e is-
ics, copy numbe al e a ions, PIK3CA mu a ion, and pa hway
signa u es. Clin Cance Res 2016;22:4735–45.
29.Ru kowski AJ, E ha d F, L’He naul A, e al. Widesp ead dis-
up ion o hos ansc ip ion e mina ion in HSV-1 in ec ion.
Na Commun 2015;6:7126.
30.Bai J, Smock SL, Jackson GR, J , e al. Pheno ypic esponses o
di e en ia ed as hma ic human ai way epi helial cul u es o
hino i us. PLoS One 2015;10:e0118286.
31.San pe e G, Da e F, Blanco S, e al. Genome-wide analysis o
wild- ype Eps ein–Ba i us genomes de i ed om heal hy
indi iduals o he 1000 Genomes P ojec . Genome Biol E ol
2014;6:846–60.
32.Mangul S, Olde Loohuis LM, O i A, e al. To al RNA sequencing
e eals mic obial communi ies in human blood and disease
speci ic e ec s , bioRxi . 2016. doi:10.1101/057570.
33.Mukhe jee S, Hun emann M, I ano a N, e al. La ge-scale
con amina ion o mic obial isola e genomes by Illumina PhiX
con ol. S and Genomic Sci 2015;10:18.
34.Wes e mann AJ, F¨
o s ne KU, Amman F, e al. Dual RNA-seq
un eils noncoding RNA unc ions in hos -pa hogen in e ac-
ions. Na u e 2016;529:496–501.
35.S ong MJ, Xu G, Mo ici L, e al. Mic obial con ami-
na ion in nex gene a ion sequencing: implica ions o
sequence-based analysis o clinical samples. PLoS Pa hog
2014;10:e1004437.
36.Bon e T, Csaba G, Zimme R, e al. Mining RNA–seq da a o
in ec ions and con amina ions. PLoS One 2013;8:e73071.
37.Cox MJ, WO C, Mo a MF. Sequencing he human mic o-
biome in heal h and disease. Hum Mol Gene 2013;22:R88–94.
38.Simon LM, Ka g S, Wes e mann A, e al. Suppo ing da a o
“Me aMap: an a las o me a ansc ip omic eads in human
disease- ela ed RNA-seq da a.” GigaScience Da abase 2018.
h p://dx.doi.o g/10.5524/100456.
40.Simon LM, Ka g S. Me aMap pipeline. p o ocols.io 2018;
doi:dx.doi.o g/10.17504/p o ocols.io.msec6be.
41 Tange O, GNU Pa allel - The Command-Line Powe ool. The
USENIX Magazine 2011;36:42–47.
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019