scieee Science in your language
[en] (orig)

MetaMap: An atlas of metatranscriptomic reads in human disease-related RNA-seq data

Abstract

Background: With the advent of the age of big data in bioinformatics, large volumes of data and high-performance computing power enable researchers to perform re-analyses of publicly available datasets at an unprecedented scale. Ever more studies imply the microbiome in both normal human physiology and a wide range of diseases. RNA sequencing technology (RNA-seq) is commonly used to infer global eukaryotic gene expression patterns under defined conditions, including human disease-related contexts; however, its generic nature also enables the detection of microbial and viral transcripts. Findings: We developed a bioinformatic pipeline to screen existing human RNA-seq datasets for the presence of microbial and viral reads by re-inspecting the non-human-mapping read fraction. We validated this approach by recapitulating outcomes from six independent, controlled infection experiments of cell line models and compared them with an alternative metatranscriptomic mapping strategy. We then applied the pipeline to close to 150 terabytes of publicly available raw RNA-seq data from more than 17,000 samples from more than 400 studies relevant to human disease using state-of-the-art high-performance computing systems. The resulting data from this large-scale re-analysis are made available in the presented MetaMap resource. Conclusions: Our results demonstrate that common human RNA-seq data, including those archived in public repositories, might contain valuable information to correlate microbial and viral detection patterns with diverse diseases. The presented MetaMap database thus provides a rich resource for hypothesis generation toward the role of the microbiome in human disease. Additionally, codes to process new datasets and perform statistical analyses are made available.

Read accessible full text

MetaMap: An atlas of metatranscriptomic reads in human disease-related RNA-seq data

Author: Simon, L. M.,Karg, S.,Westermann, A. J.,Engel, M.,Elbehery, A. H.A.,Hense, B.,Heinig, M.,Deng, L.,Theis, F. J.,Simon, L.
Publisher: Oxford University Press
Year: 2018
DOI: 10.1093/gigascience/giy070
Source: https://repository.helmholtz-hzi.de/bitstream/10033/622049/1/Simon%20et%20al.pdf
GigaScience, 7, 2018, 1–8
doi: 10.1093/gigascience/giy070
Ad ance Access Publica ion Da e: 12 June 2018
Da a No e
DATA NOTE
Me aMap: an a las o me a ansc ip omic eads in
human disease- ela ed RNA-seq da a
L.M. Simon 1,*,S.Ka g
1, A.J. Wes e mann2,3,M.Engel
1,4, A.H.A. Elbehe y5,
B. Hense1,M.Heinig
1,L.Deng
5and F.J. Theis 1,6,*
1Helmhol z Zen um M ¨
unchen, Ge man Resea ch Cen e o En i onmen al Heal h, Ins i u e o
Compu a ional Biology, Neuhe be g, Ge many, 2Ins i u e o Molecula In ec ion Biology, Uni e si y o
W¨
u zbu g, W ¨
u zbu g, Ge many, 3Helmhol z Ins i u e o RNA-Based In ec ion Resea ch, W¨
u zbu g, Ge many,
4Helmhol z Zen um M ¨
unchen, Ge man Resea ch Cen e o En i onmen al Heal h, Scien i ic Compu ing
Resea ch Uni , Neuhe be g, Ge many, 5Helmhol z Zen um M ¨
unchen, Ge man Resea ch Cen e o
En i onmen al Heal h, Ins i u e o Vi ology, Neuhe be g, Ge many and 6Depa men o Ma hema ics,
Technische Uni e si ¨
a M ¨
unchen, Munich, Ge many
∗Co espondence add ess. L.M. Simon, Helmhol z Zen um M ¨
unchen Ge man Resea ch Cen e o En i onmen al Heal h, Ins i u e o Compu a ional
Biology, Ingols ¨
ad e Lands aße, 185764, Neuhe be g; E-mail: [email p o ec ed] h p://o cid.o g/0000-0001-6148-8861; F.J. Theis;
E-mail: abian. heis@helmhol z-muenchen.de h p://o cid.o g/0000-0002-2419-1943
Abs ac
Backg ound: Wi h he ad en o he age o big da a in bioin o ma ics, la ge olumes o da a and high-pe o mance
compu ing powe enable esea che s o pe o m e-analyses o publicly a ailable da ase s a an unp eceden ed scale. E e
mo e s udies imply he mic obiome in bo h no mal human physiology and a wide ange o diseases. RNA sequencing
echnology (RNA-seq) is commonly used o in e global euka yo ic gene exp ession pa e ns unde de ined condi ions,
including human disease- ela ed con ex s; howe e , i s gene ic na u e also enables he de ec ion o mic obial and i al
ansc ip s. Findings: We de eloped a bioin o ma ic pipeline o sc een exis ing human RNA-seq da ase s o he p esence o
mic obial and i al eads by e-inspec ing he non-human-mapping ead ac ion. We alida ed his app oach by
ecapi ula ing ou comes om six independen , con olled in ec ion expe imen s o cell line models and compa ed hem
wi h an al e na i e me a ansc ip omic mapping s a egy. We hen applied he pipeline o close o 150 e aby es o publicly
a ailable aw RNA-seq da a om mo e han 17,000 samples om mo e han 400 s udies ele an o human disease using
s a e-o - he-a high-pe o mance compu ing sys ems. The esul ing da a om his la ge-scale e-analysis a e made
a ailable in he p esen ed Me aMap esou ce. Conclusions: Ou esul s demons a e ha common human RNA-seq da a,
including hose a chi ed in public eposi o ies, migh con ain aluable in o ma ion o co ela e mic obial and i al
de ec ion pa e ns wi h di e se diseases. The p esen ed Me aMap da abase hus p o ides a ich esou ce o hypo hesis
gene a ion owa d he ole o he mic obiome in human disease. Addi ionally, codes o p ocess new da ase s and pe o m
s a is ical analyses a e made a ailable.
Keywo ds: high-pe o mance compu ing; big da a; RNA-seq; sequence ead a chi e; me a ansc ip omics; mic obiome;
i ome; human disease; in ec ion
Recei ed: 22 Feb ua y 2018; Re ised: 1 June 2018; Accep ed: 4 June 2018
C
The Au ho (s) 2018. Published by Ox o d Uni e si y P ess. This is an Open Access a icle dis ibu ed unde he e ms o he C ea i e Commons
A ibu ion License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed euse, dis ibu ion, and ep oduc ion in any medium,
p o ided he o iginal wo k is p ope ly ci ed.
1
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
2Me aMap
Da a Desc ip ion
Con ex
Recen s udies ha e demons a ed he pa amoun impo ance
o he mic obiome o human heal h and disease [1]. Fo exam-
ple, imbalance o he human gu mic obiome was linked o non-
communicable diseases such as obesi y [2,3], diabe es [4], ca -
dio ascula disease [5], ch onic obs uc i e pulmona y disease
[6], and colo ec al ca cinoma [7,8], o name jus a ew.
The ad en o high- h oughpu sequencing echnologies has
e olu ionized he li e sciences. RNA sequencing (RNA-seq)
echnology p oduces one o he mos equen nex -gene a ion
sequencing da a ypes and has been applied o he s udy o a
la ge numbe o biological samples ele an o human disease.
The majo i y o he unde lying aw da a a e eely accessible
om da a eposi o ies such as he Gene Exp ession Omnibus
(>1,700 human RNA-seq da ase s as o Janua y 2018) and he Se-
quence Read A chi e (SRA) [9].
Howe e , hese da a a e ypically exclusi ely used o single
species (i.e., human) ansc ip omics such as di e en ial gene
exp ession and al e na i e splicing analysis [9,10]. Reads ha
do no map on o he human genome a e conside ed noise o
con amina ion and he e o e a e gene ally igno ed [11,12](col-
lec i ely abou 9% o o al eads, Fig. 1). Fi e yea s ago, i was pos-
ula ed ha in e species in e ac ions migh be s udied by simul-
aneous de ec ion and quan i ica ion o RNA ansc ip s om a
gi en hos and a mic obe ia “dual” RNA-seq [13]. Meanwhile,
his app oach has been success ully applied o he in e ac ion o
mammalian cells wi h di e se bac e ial [14] and i al pa hogens
[15-19].
Inspi ed by dual RNA-seq, in his s udy we hypo hesize ha
eads in a chi ed RNA-seq da ase s de i ed om human p i-
ma y cells o issue samples ha ail o map agains he hu-
man e e ence genome may con ain aluable in o ma ion abou
hep esenceo ce ainmic obesin he espec i ebodyniches
and/o unde de ined disease condi ions. To enable me a an-
sc ip omic s udy o hese da a, we combined exis ing ead align-
men and me agenomic classi ica ion so wa e in o a wo-s ep
“omni” RNA-seq pipeline o comp ehensi ely quan i y a chaeal,
bac e ial, and i al eads in human RNA-seq da a (Fig.1).
In he i s s ep o his so-called Me aMap pipeline, all eads
a e aligned agains he human genome using he ul a- as RNA-
seq aligne Spliced T ansc ip s Alignmen o a Re e ence so -
wa e (STAR) [20]. Subsequen ly, only he ac ion o unmapped
eads is subjec ed o me a ansc ip omic classi ica ion using
CLARK-S [21] (see Me hods o de ails). The combina ion be-
ween scalabili y and accu acy was he main mo i a ion behind
choosing hese wo so wa e packages o e compe ing me h-
ods [22,23]. I is impo an o no e ha CLARK-S uses a se
o uniquely disc imina i e sho sequences a he species le el
o classi y eads. The e o e, eads con aining nondisc imina i e
sequences ha ail o be uniquely assigned o a single species,
e.g., eads o igina ing om he bac e ial ibosomal 16S RNA
gene will be conside ed “unclassi ied” (al oge he 8.6% in Fig.1).
The ou pu o CLARK-S is an ope a ional axonomic uni s
(OTUs) coun ma ix, whe e ows co espond o i al, bac e ial,
and a cheal species and columns co espond o (human) sam-
ples. Each en y co esponds o he numbe o non-human eads
classi ied o he espec i e species. Fo con enience, in he ol-
lowing, we e e o he se o mic obial and i al species p o iled
using ou app oach as “me a ea u es.”
By sc eening he s udy abs ac s o he SRA o sea ch
e ms p io i izing human clinical da ase s de i ed om polyA-
independen sequencing p o ocols (see Me hods), we iden i-
ied mo e han 400 s udies ele an o human disease comp is-
ing mo e han 17,000 cDNA lib a ies (close o 150 e aby es o
aw sequencing da a). Raw sequencing eads om hese s ud-
ies we e downloaded and analyzed using he high-pe o mance
compu ing sys em o he Leibniz Supe compu ing Cen e (LRZ)
o he Ba a ian Academy o Sciences and Humani ies, which a-
cili a ed ul a- as p ocessing wi h median speeds o 25 and 21
million eads pe hou pe co e pe un o he STAR and CLARK-
S s eps, espec i ely. O he mo e han 500 billion RNA-seq eads
p ocessed, a ound 91% could be mapped o he human genome.
A ac ion o 8.6% o all eads emained nondisc imina i e a he
species le el and de ined as “unclassi ied.” In addi ion, 0.03%,
0.20%, and 0.39% o all eads we e assigned o a chaeal, bac e-
ial, o i al me a ea u es, espec i ely. Despi e hese ela i ely
low pe cen ages, he absolu e numbe s o eads classi ied we e
in he hund ed millions o billions, enabling s a is ical analyses.
Me hods
High-pe o mance compu ing en i onmen
P ojec compu a ions including download, alignmen o eads
on o he human genome, and me a ea u e quan i ica ion we e
made on he high-pe o mance Linux Clus e a he LRZ [24].
RNA-seq da a e ie al
Raw nex -gene a ion sequencing da a we e downloaded om
he SRA. The R package SRAdb was downloaded on 23 May 2017
and used o que y he SRA da abase. To iden i y SRA p ojec s
ha con ain ansc ip omic analyses o human RNA-seq da a,
he SRA a ibu es “ axon id,” “lib a y sou ce,” “lib a y s a egy,”
and “pla o m” we e sea ched o he e ms “9606,” “TRAN-
SCRIPT,” “RNA-seq,” and “ILLUMINA,” espec i ely. To emo e
po en ial bias de i ed om di e en sequencing echnologies,
we also es ic ed he que y o SRA uns anno a ed wi h “ILLU-
MINA” in SRA a ibu e “pla o m.” To exclude s udies wi h in-
su icien sample size o s a is ical analysis, he que y was e-
s ic ed o SRA p ojec s con aining mo e han i e uns. To a oid
concen a ing he analysis on a small numbe o la ge p ojec s,
he que y was es ic ed o SRA p ojec s wi h ewe han 500
uns. To iden i y s udies ocusing on pheno ypes ele an o hu-
man disease, we es ic ed he que y o uns con aining a leas
one o mo e o he e ms “disease,” “pa ien ,” “p ima y,” and
“clinical” in he SRA a ibu e “s udy abs ac .” To exclude in i o
(cell-cul u e) expe imen s bu ocus on p ima y (clinical) sam-
ples, SRA uns con aining he e ms “mu an ” o “cell-line” we e
emo ed om ou selec ion. Fu he mo e, SRA uns con aining
he e ms “single cell” and “GTEx” we e emo ed. Finally, sam-
ples wi h ewe han 1 million o al eads o ead leng hs <50 bp
we e excluded. The desc ibed que y esul ed in 484 sho ead
p ojec s (SRPs) con aining 21,659 RNA-seq uns. Due o echnical
p oblems (i.e., missing URLs, es ic ed access), we we e unable
o download a ac ion o 4,078 samples.
Human alignmen
Alignmen o eads agains he human e e ence genome (hg38)
and simul aneous human gene exp ession quan i ica ion was
conduc ed wi h STAR ( e sion 2.5.2). To inc ease mapping speed
o a la ge numbe o samples, we used he –genomeLoad LoadAnd-
Keep unc ion o load he STAR index once and keep i in
memo y o subsequen alignmen s. The pa ame e –quan mode
GeneCoun s was used o gene a e he human gene exp es-
sion coun ables. Unmapped eads we e sa ed wi h he –
ou ReadsUnmapped Fas x pa ame e . To u he inc ease mapping
speed, mul iple h eads we e used as implemen ed wi h he pa-
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
Simon e al. 3
Figu e 1: Schema ic o he Me aMap pipeline. Mo e han 400 p ojec s om s udies ele an o human disease we e iden i ied in he SRA da abase. Mo e han 500 billion
RNA-seq eads we e downloaded and i s il e ed by mapping hem on o he human genome. The emaining eads unde wen me a ea u e classi ica ion. I is no ed
ha 90.7% o all eads mapped o he human genome; 0.03%, 0.20%, and 0.39% o all eads we e assigned o a chaeal, bac e ial, o i al me a ea u es, espec i ely; and
8.6% o all eads emain nondisc imina i e a he species le el (“unclassi ied”).
ame e – unTh eadN 28. Runs wi h ewe han 30% eads map-
ping o he human genome we e excluded om downs eam
analysis. All human alignmen s we e conduc ed on he LRZ
“CoolMUC2” Linux-Clus e . This clus e con ains 384 nodes wi h
64 GB andom access memo y (RAM) memo y and 28 co es each.
Me a ea u e quan i ica ion
Me a ea u e quan i ica ion was conduc ed wi h CLARK-S ( e -
sion 1.2.3). CLARK-S is a so wa e me hod o as and accu-
a e sequence classi ica ion o me agenomic nex -gene a ion
sequencing da a, including RNA-seq da a. One majo issue du -
ing he classi ica ion o me agenomic da a is he ising numbe
o a ge s o align agains . CLARK-S sol es his issue by build-
ing a la ge index ile consis ing o disc imina i e k-me s. The
me agenomic e e ence da abase was gene a ed ollowing he
desc ip ion o he CLARK websi e using he ollowing wo com-
mands: se a ge s.sh bac e ia i us –species and buildSpacedDB.sh.
This da abase con ained 16,551 genome sequences co espond-
ing o 6,979 unique species (Addi ional File 2). To allow uni-
o m p ocessing, pai ed-end sequencing expe imen s we e an-
alyzed independen ly. Each single unmapped ead ile was
used as inpu o CLARK-S wi h he ollowing pa ame e s: clas-
si y me agenome.sh –spaced –O lis o FASTQ iles. To inc ease
classi ica ion speed, he CLARK-S exp ess mode was selec ed
and mul iple h eads we e used wi h pa ame e s –m 2 and –n
32, espec i ely. The ou pu iles o his s ep con ain all inpu
ead iden i ie s wi h he co esponding me a ea u e classi ica-
ion. In he subsequen s ep, o al coun s a e summa ized o
each ea u e wi h he es ima e abundance.sh command. To en-
able compa ison ac oss single-end and pai ed-end expe imen s,
me a ea u e coun s om pai ed-end expe imen s we e a e -
aged and subsequen ly ounded o conse e coun dis ibu ion.
To accoun o a ying sequencing dep hs, me a ea u e abun-
dance was es ima ed as he numbe o eads pe million o-
al eads sequenced. Me a ea u e quan i ica ion was conduc ed
on he LRZ “Te amem” Linux-Clus e . This clus e con ains one
node wi h 6,144 GB RAM memo y and 96 co es.
BLAST-based me a ea u e classi ica ion
To alida e esul s gene a ed by he Me aMap pipeline, he Basic
Local Alignmen Sea ch Tool (BLAST) [25] was used as ollows. A
BLAST da abase was c ea ed om he same genome sequences
used in he CLARK-S app oach. Then, eads we e aligned o
his da abase using BLASTN wi h a h eshold E- alue o 1e-10.
P oduced coun s om pai ed-end expe imen s we e a e aged.
Fo each ile, BLAST was done by unning app oxima ely 10 kb
chunks ( eco d sepa a o “>”) in pa allel using pa allel [41] (28
jobs), each wi h eigh h eads using one node on he LRZ “Cool-
MUC3” Linux Clus e . This clus e con ains 148 nodes wi h 96
GB RAM memo y and 64 co es each. Ou pu was pa sed o ex-
clusi ely keep eads ha could be assigned a he species le el.
Di e en ial me a ea u e abundance
Di e en ial me a ea u e abundance analysis was pe o med us-
ing he R package DESeq2 [26]. DESeq2 models di e en ial gene
exp ession by i ing a nega i e binomial dis ibu ion o he aw
coun s unde lying RNA-seq da a. This amewo k can accoun
o con ounding a iables such as sequencing dep h. The e o e,
he da a need no be no malized p io o s a is ical in e -sample
compa isons. Fo each o he ou published bona ide dual RNA-
seq s udies, we classi ied samples in o he ollowing wo g oups
based on he p o ided anno a ions: samples expec ed o con ain
he known pa hogen, such as human papilloma i us-posi i e
umo s in he Zhang e al. s udy [28], and pa hogen- ee con-
ols, such as mock- ea ed cells in he Wes e mann e al. [27]
s udy. Using his bina y ou come, we pe o med di e en ial ex-
p ession analysis ac oss all de ec ed me a ea u es. To accoun
o sequencing dep h, lib a y size ac o s we e es ima ed om
he o al numbe o sequenced eads. The dispe sion o he neg-
a i e binomial dis ibu ion was es ima ed using a local linea e-
g ession as implemen ed in he DESeq() unc ion ia he i Type
pa ame e “local.”
Da a alida ion and quali y con ol
We alida ed ou app oach by eco e ing he g ound u h in
bona ide dual RNA-seq expe imen s pe o med wi h human
cell lines and samples om pa ien s wi h well-known in ec-
ion s a us. O he ou selec ed s udies, one analyzed an in ec-
ion model based on a bac e ial (Salmonella en e ica se o a Ty-
phimu ium) and h ee based on dis inc i al pa hogens (human
papilloma i us, he pes simplex i us, hino i us). As expec ed,
Me aMap de ec ed he known pa hogen a highe le els in he
espec i e s udy compa ed o he o he s udies and pa hogens
(Table1). Howe e , compa isons ac oss s udies and me a ea u es
may be biased by echnical con ounde s (discussed in de ail in
he Re-use po en ial sec ion). The e o e, we ocused ou analy-
sis on he compa ison o a single me a ea u e ac oss subjec s
wi hin a s udy. Using he anno a ion p o ided in he espec i e
s udy, we pe o med di e en ial me a ea u e abundance analy-
sis o iden i y hose me a ea u es ha show he la ges ela i e
di e ence in abundance le els be ween he in ec ed and con ol
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
4Me aMap
Table 1: O e iew o ou dual RNA-seq s udies used o alida e he Me aMap pipeline.
S udy
In ec ion
agen
To al
eads
Salmonella
en e ica
Alphapapilloma i us
9
Human
alphahe pes i us
1
Rhino i us
A
Wes e mann
e al. [27]
Salmonella en e ica se o a
Typhimu ium
1.0e+07 6.3e+03 1.2e-01 1.5e-01 1.2e-01
Zhang e al.
[28]
Human papilloma i us 4.6e+07 3.0e-02 5.1e+01 2.2e-02 2.2e-02
Ru kowski e
al. [29]
He pes simplex i us 3.5e+07 1.1e+00 3.1e-02 3.1e+04 3.0e-02
Bai e al. [30] Rhino i us 6.6e+06 2.0e-01 1.5e-01 1.5e-01 4.4e+01
To al eads column depic s he a e age ead dep h pe sample o each s udy. A e age me a ea u e abundance o alphapapilloma i us 9, Salmonella en e ica, human
alphahe pes i us 1, and hino i us A a e shown in eads pe million. The co ec in ec ion agen o he espec i e s udy is highligh ed in bold on
0
25
50
75
0510
−log10 p− alue
Wes e mann e al
0
5000
10000
15000
in ec ed
(n=36)
mock
(n=6)
Salmonella en e ica
A
0
10
20
04812
Zhang e al
Alphapapilloma i us 9
B
0.0
5.0
10.0
−5.0 −2.5 0.0 2.5 5.0 7.5
Ru kowski e al
0
25000
50000
75000
in ec ed
(n=8)
mock
(n=2)
H. Alphahe pes i us 1
C
0
20
40
048
Fold change (log2)
Bai e al
0
50
100
150
200
in ec ed
(n=12)
ehicle
(n=12)
Rhino i us A
D
0
100
200
300
400
500
nega i e
(n=18)
posi i e
(n=18)
Figu e 2: Di e en ial me a ea u e abundance analysis o con olled in ec ion expe imen s eco e s g ound u h. “Volcano” plo s show old change and in e ed P alue
on he xand yaxes, espec i ely. Each do ep esen s a me a ea u e. The mos signi ican me a ea u e is colo ed in ed. Inse s display box plo s o he abundance
le els in eads pe million o he op hi me a ea u e ac oss condi ions o each s udy. Fo all box plo s, he box ep esen s he in e qua ile ange, he ho izon al line
in he box is he median, and he whiske s ep esen 1.5 imes he in e qua ile ange.
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
Simon e al. 5
Figu e 3: Analysis o lymphoblas cell line expe imen s u he suppo s he Me aMap pipeline. (A and B) Mean abundance le els ac oss all samples o he op i e
me a ea u es o p ojec s SRP041338 and SRP091453, espec i ely. (C) Rela i e p opo ion o eads mapping o EBV, phiX, and all o he me a ea u es ac oss RNA-seq
samples. (D) Cumula i e dis ibu ion plo o he a e age p opo ion o bac e ial me a ea u e eads ac oss all p ojec s. Pu ple and pink e ical lines highligh p ojec s
SRP041338 and SRP091453, espec i ely.
samples (see Me hods o de ails). The co ec in ec ion agen
showed he mos signi ican di e ence ac oss all me a ea u es
be ween in ec ed and con ol samples o each s udy (Fig.2). Fo
example, Wes e mann e al. [27] gene a ed dual RNA-seq da a
om HeLa cells in ec ed wi h he en e ic bac e ial pa hogen
S. en e ica se o a Typhimu ium and compa ed hem o mock-
ea ed con ol samples. Acco dingly, we obse ed S. en e ica as
he mos di e en ially abundan me a ea u e be ween he in-
ec ed and he con ol samples (P<1e-75, Fig. 2A). Likewise, we
eco e ed alphapapilloma i us 9,human alphahe pes i us 1 (also
known as he pes simplex i us 1), and hino i us A as he mos
di e en ially abundan me a ea u es in he da a om Zhang e
al. [28], Ru kowski e al. [29], and Bai e al. [30], espec i ely. In he
Wes e mann e al. [27] and Ru kowski e al. [29] s udies, se e al
addi ional me a ea u es showed a s ong di e en ial abundance
e ec (Fig. 2A and 2C). These me a ea u es we e closely ela ed
o he ue in ec ion agen , i.e., Salmonella bongo i (P<1e-67) and
Panine alphahe pes i us 3 (P<1e-9) o he Wes e mann e al. [27]
o Ru kowski e al. [29] s udy, espec i ely. These indings con-
i m ha ou Me aMap pipeline ecapi ula es esul s om dedi-
ca ed dual RNA-seq s udies, i.e., s udies based on known in ec-
ious agen s. The e o e, Me aMap may be equally sui ed o de-
ec p e iously unknown mic obial and i al species in human
p ima y samples.
As an addi ional con ol, we e-analyzed wo p ojec s con-
ained in ou da a collec ion ha a e de i ed om he B lym-
phoblas cell line unde nonin ec ious condi ions. Howe e ,
since Eps ein-Ba i us (EBV) is used o ans ec ion and ans-
o ma ion o lymphocy es o lymphoblas s, we expec ed o de-
ec eads om his i us in hese p ojec s [31], bu no u -
he i al o mic obial eads [32]. Indeed, he mos abundan
me a ea u es in each p ojec we e domina ed by eads classi-
ied o gammahe pes i us 4 (also known as EBV) and En e obac e-
ia phage phiX174 sensu la o (phiX), commonly used as spike-in
in Illumina sequencing uns [33](Fig.3A and 3B). On a e age,
95% and 97% o all me a ea u e eads we e classi ied as phiX o
EBV o p ojec s SRP041338 and SRP091453, espec i ely (Fig. 3C).
Con e sely, he abundance o eads mapping o bac e ial species
o hese wo p ojec s co esponds o he bo om pe cen ile as
compa ed o all o he p ojec s in he Me aMap da abase, sup-
po ing s e ili y o his cell line (Fig. 3D). This demons a es ha
Me aMap no only is capable o edisco e ing known pa hogenic
species ( ue posi i es) in con olled in ec ion expe imen s (Fig.
2) bu i also minimizes he de ec ion o alse posi i es o , a
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019

6Me aMap
0 5000 10000 15000
0 5000 10000 15000
Salmonella en e ica le els
CLARK-S based abundance (RPM)
BLAST based abundance (RPM)
Spea man
Rho 0.98
P < 2.2e-16
1 5 50 500 5000
1 10 100 1000 10000
A e age me a ea u e abundance
CLARK-S based (RPM)
BLAST based (RPM)
Spea man
Rho 0.16
P 3.1e−10
AB
C
Salmonella en e ica
CLARK−S
BLAST
0246
Reads pe hou pe h ead
(log10)
8
Figu e 4: Al e na i e BLAST-based classi ica ion me hod alida es me a ea u e abundance es ima es by Me aMap. (A) A e age me a ea u e eads pe million le els
de i ed using he CLARK-S so wa e, as implemen ed in he Me aMap pipeline, and a BLAST-based al e na i e app oach on he xand yaxes, espec i ely. (B) Co ela ion
in S. en e ica abundance le els be ween he wo classi ica ion app oaches. (C) Di e ence in classi ica ion speed be ween he BLAST and CLARK-S me a ansc ip omic
classi ica ion. The yaxis shows he numbe o eads p ocessed pe hou pe h ead in log10 space.
leas , p o ides measu es such as abundance and signi icance,
allowing he use o iden i y and coun e selec hose species.
As a echnical alida ion, we compa ed ou app oach o an
al e na i e me a ansc ip omic classi ica ion s a egy o he
Wes e mann e al. [34] s udy. All non-human eads we e aligned
using BLASTN o a BLAST da abase consis ing o he same ge-
nomic sequences used by CLARK-S (see Me hods o de ails). The
a e age me a ea u e abundances ac oss all 42 samples de i ed
om he BLAST-based app oach and CLARK-S co ela ed sig-
ni ican ly (Spea man co ela ion, Rho: 0.16, P: 3.1e-10) (Fig. 4A).
BLAST showed highe sensi i i y and de ec ed mo e me a ea-
u es compa ed o CLARK-S (indica ed by he accumula ion o
do s a alue 0 on he xaxis in Fig. 4A).Thisismos lyobse ed
o low abundance me a ea u es ha could ep esen low coun s
de i ed om sequencing and/o mapping e o s. Howe e , mos
impo an ly, he ue pa hogen me a ea u e “Salmonella en e -
ica” showed e y high co ela ion ac oss samples be ween he
BLAST- and CLARK-based abundance es ima es (Fig. 4B). No e-
wo hy, he Me aMap pipeline p ocessed eads mo e han h ee
o de s o magni ude as e han BLAST, demons a ing a sig-
ni ican speed ad an age while gene a ing compa able esul s
(Fig. 4C).
Re-use po en ial
Mic obial and i al con amina ion in nex -gene a ion sequenc-
ing da a has been obse ed. I can be caused by mapping e o s
due o genome sequence simila i y be ween di e en species
[35,36]. In addi ion, echnical con ounde s can obs uc he
analysis and po en ially gene a e a i icial di e ences i no con-
side ed p ope ly. Fo example, di e en ypes o human sam-
ples may con ain di e en amoun s o non-human ma e ial
due o a ying s e ili y o he issues. Fu he mo e, sequencing
dep h may in oduce a de ec ion loo o me a ea u es ha a e
no abundan . The e o e, compa isons ac oss di e en issues
and sequencing dep hs may gene a e a i icial di e ences. Addi-
ionally, gi en ha only uniquely disc imina i e sequences a e
coun ed, he absolu e abundance le els may no be compa a-
ble ac oss me a ea u es. Finally, he Me aMap pipeline cap u es
me a ea u e abundance a he RNA le el, which may no nec-
essa ily co espond o genomic abundance le els. Me a ea u es
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
Simon e al. 7
may no be abundan a he DNA le el bu highly ansc ip ion-
ally ac i e and hus abundan ly de ec ed a he RNA le el, o he
in e se. These po en ial challenges need o be aken in o con-
side a ion when compa ing ac oss me a ea u es.
To minimize hese e ec s, we encou age ocusing on s udies
ha include in ap ojec compa isons ha es one me a ea u e
a a ime, as exempli ied in he di e en ial me a ea u e abun-
dance analysis. Ou a ionale is ha echnical con ounde s, in
con as o biologically meaning ul changes, should a ec all
uns wi hin a p ojec o he same ex en and he e o e no show
condi ion-speci ic e ec s. Fo example, in he Wes e mann e al.
s udy [34], we de ec ed subs an ial le els o phiX in bo h condi-
ions (in ec ed samples and mock- ea ed con ols), bu only he
“Salmonella” me a ea u e showed a condi ion-speci ic e ec . We
aim o add ess he challenges inhe en o in e p ojec and in e -
me a ea u e compa isons in u u e wo k.
All he aw da a desc ibed in he p esen s udy we e pub-
licly a ailable, ye ha e been e y cumbe some o ex ac in-
di idually. The p esen ed Me aMap da abase makes hese da a
easily accessible o a e y b oad communi y, he eby allow-
ing o global compa isons o e hund eds o indi idual s udies
and housands o sampled condi ions. While we a emp ed o
minimize he isk o de ec ing alse posi i es (Fig. 3), i should
be no ed ha no all me a ea u es classi ied by Me aMap will
necessa ily e e o ue biological ac o s. No ewo hy, ou ap-
p oach e eals a co ela ion be ween me a ea u es and disease,
no causali y, and canno disc imina e disease-associa ed e -
ec s om po en ial ea men e ec s. Howe e , ou pipeline
p o ides he use wi h a scien i ic s a ing g ound o alida e he
p esence/absence o de ined mic obial and i al species unde
de ined condi ions and explo e he unde lying biology and sig-
ni icance in g ea e de ail. As a po en ial use case o hese da a,
use s can es o associa ions o mic obial o i al me a ea-
u es wi h a ple ho a o human diseases o be ween hemsel es.
In addi ion, use s wi h in e es in a speci ic bac e ial o i al
species can easily iden i y s udies and, consequen ly, disease
con ex s in which eads om his o ganism we e de ec ed. This
could gi e an impo an i s hin o assess whe he he espec-
i e species migh be implica ed in a gi en human disease e i-
ology. Fu he mo e, his esou ce p o ides he oppo uni y o
suppo indings de i ed om s anda d mic obiome p o iling
echnologies, such as 16S RNA gene based o sho gun me age-
nomics [37]. Finally, me a ea u e de ec ion in human clinical
RNA-seq samples may p o ide a diagnos ic ad an age when
s udying mic obes o i uses ha a e challenging o isola e.
The composi e me a ea u e OTU coun able, de i ed om
17 278 cDNA lib a ies om 436 SRA p ojec s, including anno a-
ions is p o ided o download [38].
A ailabili y o sou ce code and equi emen s
P ojec name: Me aMap
P ojec home: h ps://gi hub.com/ heislab/Me aMap
Ope a ing sys em(s): Pla o m-independen
P og amming language: Unix command line, R
O he equi emen s: STAR and CLARK-S may equi e la ge
amoun s o memo y (>100 GB)
License: GNU GPL
A ailabili y o suppo ing da a
The da ase s suppo ing he esul s p esen ed he e a e a ailable
in he GigaScience Da abase eposi o y [38]. The p o ocols a e also
a ailable a [40].
Addi ional ile
Addi ional File1.cs
Abb e ia ions
BLAST: basic local alignmen sea ch ool; EBV: Eps ein-Ba i us;
LRZ: Leibniz Supe compu ing Cen e; OTU: ope a ional axo-
nomic uni ; phiX: En e obac e ia phage phiX174 sensu la o; RNA-
seq: RNA sequencing; SRA: Sequencing Read A chi e; SRP: sho
ead p ojec ; STAR: Spliced T ansc ip s Alignmen o a Re e ence
so wa e.
Compe ing in e es s
The au ho s decla e ha hey ha e no compe ing in e es s.
Funding
L.S. acknowledges unding om he Eu opean Union’s Ho i-
zon 2020 Resea ch and Inno a ion P og amme unde he Ma ie
Sklodowska-Cu ie g an ag eemen (753039). The ope a ion o
he LRZ Linux Clus e is unded ia he Ba a ian S a e Minis y
o Educa ion, Science, and he A s.
Au ho con ibu ions
Concep ualiza ion: L.S., M.E., L.D., and B.H.; o mal analysis: L.S.,
M.H., S.K., and A.E.; in es iga ion: L.S., A.J.W., and M.E.; me hod-
ology: L.S., S.K, M.H.; w i ing he o iginal d a : L.S. and A.J.W;
w i ing, e iewing, and edi ing: L.S., A.J.W., M.E., A.E., L.D., M.H.,
and F.T.; supe ision: L.D., M.H., and F.T.
Acknowledgmen s
The au ho s hank Yu Wang and Fe dinand Jami zky om he
LRZ o hei suppo .
Re e ences
1. Young VB. The ole o he mic obiome in human heal h and
disease: an in oduc ion o clinicians. BMJ 2017;356:j831.
2. Tu nbaugh PJ, Ley RE, Mahowald MA, e al. An obesi y-
associa ed gu mic obiome wi h inc eased capaci y o en-
e gy ha es . Na u e 2006;444:1027–31.
3. Henao-Mejia J, Elina E, Jin C, e al. In lammasome-media ed
dysbiosis egula es p og ession o NAFLD and obesi y. Na u e
2012;482:179–85.
4. Cani PD, Bibiloni R, Knau C, e al. Changes in gu mic o-
bio a con ol me abolic endo oxemia-induced in lamma ion
in high- a die -induced obesi y and diabe es in mice. Dia-
be es 2008;57:1470–81.
5. Wang Z, Klip ell E, Benne BJ, e al. Gu lo a me abolism
o phospha idylcholine p omo es ca dio ascula disease. Na-
u e 2011;472:57–63.
6. Engel M, Endes elde D, Schlo e -Hai B, e al. In luence o lung
CT changes in ch onic obs uc i e pulmona y disease (COPD)
on he human lung mic obiome. PLoS One 2017;12:e0180859.
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019
8Me aMap
7. Kos ic AD, Ge e s D, Pedamallu CS, e al. Genomic analysis
iden i ies associa ion o Fusobac e ium wi h colo ec al ca ci-
noma. Genome Res 2012;22:292–8.
8. Cas ella in M, Wa en RL, F eeman JD, e al. Fusobac e ium nu-
clea um in ec ion is p e alen in human colo ec al ca cinoma.
Genome Res 2012;22:299–306.
9. Kodama Y, Shumway M, Leinonen R, In e na ional Nu-
cleo ide Sequence Da abase Collabo a ion. The Sequence
Read A chi e: explosi e g ow h o sequencing da a. Nucleic
Acids Res 2012;40:D54–6.
10.Conesa A, Mad igal P, Ta azona S, e al. A su ey o bes p ac-
ices o RNA-seq da a analysis. Genome Biol 2016;17:13.
11.Gouin A, Legeai F, Nouhaud P, e al. Whole-genome e-
sequencing o non-model o ganisms: lessons om un-
mapped eads. He edi y 2015;114:494–501.
12.Peng X, Wang J, Zhang Z, e al. Re-alignmen o he un-
mapped eads wi h base quali y sco e. BMC Bioin o ma ics
2015;16(Suppl 5):S8.
13.Wes e mann AJ, Go ski SA, Vogel J. Dual RNA-seq o pa hogen
and hos . Na Re Mic obiol 2012;10:618–30.
14.Wes e mann AJ, Ba quis L, Vogel J. Resol ing hos -pa hogen
in e ac ions by dual RNA-seq. PLoS Pa hog 2017;13:e1006033.
15.Ju anic Lisnic V, Babic Cac M, Lisnic B, e al. Dual analysis o
he mu ine cy omegalo i us and hos cell ansc ip omes e-
eal new aspec s o he i us-hos cell in e ace. PLoS Pa hog
2013;9:e1003611.
16.Xu G, S ong MJ, Lacey MR, e al. RNA CoMPASS: a dual ap-
p oach o pa hogen and hos ansc ip ome analysis o RNA-
seq da ase s. PLoS One 2014;9:e89445.
17.Pa k S-J, Kuma M, Kwon H-I, e al. Dynamic changes in hos
gene exp ession associa ed wi h H5N8 a ian in luenza i us
in ec ion in mice. Sci Rep 2015;5:16512.
18.Saxena K, Simon LM, Zeng X-L, e al. A pa adox o ansc ip-
ional and unc ional inna e in e e on esponses o human
in es inal en e oids o en e ic i us in ec ion. P oc Na l Acad
Sci 2017;114:E570–9.
19.Wesolowska-Ande sen A, E e man JL, Da idson R, e al. Dual
RNA-seq e eals i al in ec ions in as hma ic child en wi h-
ou espi a o y illness which a e associa ed wi h changes in
he ai way ansc ip ome. Genome Biol 2017;18:12.
20.Dobin A, Da is CA, Schlesinge F, e al. STAR: ul a as uni e -
sal RNA-seq aligne . Bioin o ma ics 2012;29:15–21.
21.Ouni R, Lona di S. Highe classi ica ion sensi i i y o
sho me agenomic eads wi h CLARK-S. Bioin o ma ics
2016;32:3823–5.
22.Lindg een S, Adai KL, Ga dne PP. An e alua ion o he ac-
cu acy and speed o me agenome analysis ools. Sci Rep
2016;6:19233.
23.Engs ¨
om PG, S eijge T, Sipos B, e al. Sys ema ic e alua ion
o spliced alignmen p og ams o RNA-seq da a. Na Me h-
ods 2013;10:1185–91.
24. www.l z.de/se ices/compu e/linux-clus e , Leibniz Supe -
compu ing Cen e
25.Al schul S. Basic local alignmen sea ch ool. J Mol Biol
1990;215:403–10.
26.Lo e MI, Hube W, Ande s S. Mode a ed es ima ion o old
change and dispe sion o RNA-seq da a wi h DESeq2 [In e -
ne ]. Genome Biol 2014;15:550.
27.Wes e mann AJ, F¨
o s ne KU, Amman F, e al. Dual RNA-seq
un eils noncoding RNA unc ions in hos –pa hogen in e ac-
ions. Na u e 2016;529:496–501.
28.Zhang Y, Kone a LA, Vi ani S, e al. Sub ypes o HPV-posi i e
head and neck cance s a e associa ed wi h HPV cha ac e is-
ics, copy numbe al e a ions, PIK3CA mu a ion, and pa hway
signa u es. Clin Cance Res 2016;22:4735–45.
29.Ru kowski AJ, E ha d F, L’He naul A, e al. Widesp ead dis-
up ion o hos ansc ip ion e mina ion in HSV-1 in ec ion.
Na Commun 2015;6:7126.
30.Bai J, Smock SL, Jackson GR, J , e al. Pheno ypic esponses o
di e en ia ed as hma ic human ai way epi helial cul u es o
hino i us. PLoS One 2015;10:e0118286.
31.San pe e G, Da e F, Blanco S, e al. Genome-wide analysis o
wild- ype Eps ein–Ba i us genomes de i ed om heal hy
indi iduals o he 1000 Genomes P ojec . Genome Biol E ol
2014;6:846–60.
32.Mangul S, Olde Loohuis LM, O i A, e al. To al RNA sequencing
e eals mic obial communi ies in human blood and disease
speci ic e ec s , bioRxi . 2016. doi:10.1101/057570.
33.Mukhe jee S, Hun emann M, I ano a N, e al. La ge-scale
con amina ion o mic obial isola e genomes by Illumina PhiX
con ol. S and Genomic Sci 2015;10:18.
34.Wes e mann AJ, F¨
o s ne KU, Amman F, e al. Dual RNA-seq
un eils noncoding RNA unc ions in hos -pa hogen in e ac-
ions. Na u e 2016;529:496–501.
35.S ong MJ, Xu G, Mo ici L, e al. Mic obial con ami-
na ion in nex gene a ion sequencing: implica ions o
sequence-based analysis o clinical samples. PLoS Pa hog
2014;10:e1004437.
36.Bon e T, Csaba G, Zimme R, e al. Mining RNA–seq da a o
in ec ions and con amina ions. PLoS One 2013;8:e73071.
37.Cox MJ, WO C, Mo a MF. Sequencing he human mic o-
biome in heal h and disease. Hum Mol Gene 2013;22:R88–94.
38.Simon LM, Ka g S, Wes e mann A, e al. Suppo ing da a o
“Me aMap: an a las o me a ansc ip omic eads in human
disease- ela ed RNA-seq da a.” GigaScience Da abase 2018.
h p://dx.doi.o g/10.5524/100456.
40.Simon LM, Ka g S. Me aMap pipeline. p o ocols.io 2018;
doi:dx.doi.o g/10.17504/p o ocols.io.msec6be.
41 Tange O, GNU Pa allel - The Command-Line Powe ool. The
USENIX Magazine 2011;36:42–47.
Downloaded om h ps://academic.oup.com/gigascience/a icle-abs ac /7/6/giy070/5036539 by Ges Bio echnologische use on 11 Decembe 2019