scieee Science in your language
[en] (orig)

Vipie: web pipeline for parallel characterization of viral populations from multiple NGS samples

Abstract

BioMed Central open access

Read accessible full text

Vipie: web pipeline for parallel characterization of viral populations from multiple NGS samples

Author: Lin, Jake,Kramka, Lenka,Autio, Reija,Hyöty, Heikki,Nykter, Matti,Cinek, Ondrej
Year: 2017
Source: https://trepo.tuni.fi/bitstream/10024/101314/1/vipie_web_pipeline_2017.pdf
SOFTWARE Open Access
Vipie: web pipeline o pa allel
cha ac e iza ion o i al popula ions
om mul iple NGS samples
Jake Lin
1
, Lenka K amna
2
, Reija Au io
3
, Heikki Hyö y
1,4*
, Ma i Nyk e
1*
and Ond ej Cinek
2*
Abs ac
Backg ound: Nex gene a ion sequencing (NGS) echnology allows labo a o ies o in es iga e i ome composi ion
in clinical and en i onmen al samples in a cul u e-independen way. The e is a need o bioin o ma ic ools capable
o pa allel p ocessing o i ome sequencing da a by exac ly iden ical me hods: his is especially impo an in s udies
o mul i ac o ial diseases, o in pa allel compa ison o labo a o y p o ocols.
Resul s: We ha e de eloped a web-based applica ion allowing di ec upload o sequences om mul iple i ome
samples using cus om pa ame e s. The samples a e hen p ocessed in pa allel using an iden ical p o ocol, and can
be easily eanalyzed. The pipeline pe o ms de-no o assembly, axonomic classi ica ion o i uses as well as sample
analyses based on use -de ined g ouping ca ego ies. Tables o i us abundance a e p oduced om c oss- alida ion
by emapping he sequencing eads o a union o all obse ed e e ence i uses. In addi ion, ead se s and epo s
a e c ea ed a e p ocessing unmapped eads agains known human and bac e ial ibosome e e ences. Secu ed
in e ac i e esul s a e dynamically plo ed wi h popula ion and di e si y cha s, clus e ed hea maps and a so able
and sea chable abundance able.
Conclusions: The Vipie web applica ion is a unique ool o mul i-sample me agenomic analysis o i al da a,
p oducing sea chable hi s ables, in e ac i e popula ion maps, alpha di e si y measu es and clus e ed hea maps
ha a e g ouped in applicable cus om sample ca ego ies. Known e e ences such as human genome and bac e ial
ibosomal genes a e op ionally emo ed om unmapped (‘da k ma e ’) eads. Secu ed esul s a e accessible and
sha eable on mode n b owse s. Vipie is a eely a ailable web-based ool whose code is open sou ce.
Keywo ds: Me agenomics, Vi omes, Vi us, Assembly, NGS analysis, Visualiza ion, Pa allel p ocessing, Vi al da k ma e
Backg ound
The use o i ome me agenomics has been g owing
apidly due o he inc easing demands o s udy he
whole i ome in clinical samples and o e alua e he
e olu ion o i al quasispecies du ing acu e and ch onic
in ec ions. The applica ion o i ome sequencing ech-
niques become use ul no only in in ec ious disease
esea ch, bu also in associa ion s udies o p ima ily
non-in ec ious condi ions, i.e. in diseases whe e he
agen is p esumed o modi y he isk o he disease,
which e ec is de ec able upon in es iga ion o a la ge
numbe o subjec s only. These applica ions equi e an
app oxima ion o i us quan i y, simila o wha has long
been u ilized in bac e iome p o iling.
As i uses lack a common sequence signa u e, me age-
nomics sequencing o andom i al lib a ies emains he
only easible way o an unbiased assessmen o he whole
i ome. P esen ly, he need o accu a e quan i ica ion
and in e p e a ion o i al popula ion me ics ac oss a
se o samples c ea es a subs an ial challenge o his
kind o me agenomics s udies. P ime obs acles o
i ome in es iga o s a e he la ge gene ic he e ogenei y
and also ha he majo i y o bioin o ma ic ools a e
command line based and o e ly echnical, being com-
pu a ionally demanding, wi h complica ed dependencies,
* Co espondence: [email p o ec ed];ma i.nyk e[email p o ec ed];
[email p o ec ed]
1
BioMediTech and Facul y o Medicine and Li e Sciences, Uni e si y o
Tampe e, PB 100FI-33014 Tampe e, Finland
2
Depa men o Pedia ics, 2nd Facul y o Medicine, Cha les Uni e si y and
Uni e si y Hospi al Mo ol, V Ú alu 84, 150 06 P aha 5, Czech Republic
Full lis o au ho in o ma ion is a ailable a he end o he a icle
© The Au ho (s). 2017 Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0
In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and
ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o
he C ea i e Commons license, and indica e i changes we e made. The C ea i e Commons Public Domain Dedica ion wai e
(h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed.
Lin e al. BMC Genomics (2017) 18:378
DOI 10.1186/s12864-017-3721-7
and p oducing ex based ou pu s ha a e no easily
in e p e able [1–5]. Recen ly eleased web based applica-
ions Taxonome [6], Vi usTAP [7], Vi ome [8] and
Me a i [9, 10] ha e add essed some o he issues (espe-
cially hose o use in e ac ion), bu mos ly ope a e only
on single sample expe imen s wi h di e en wo k lows.
Requi ing local dependencies and ins alla ion, Vi omeScan
[11] and Me aSho [12] wo ks on mul iple samples. Some
o hese ools we e designed o long (>300) eads o
assembled con igs [8–10], which is limi ing as mode n
me agenomics p ojec s including Human Mic obiome
P ojec (HMP) [1, 2] p oduce mos ly high- h oughpu
sho pai ed eads. Table 1 p o ides an o e iew o he
p ima y ea u es and s a egies o hese di e en ools,
including ou wo k.
We aimed o open he possibili y o c ea ing a able o
i al quan i ies o mul iple samples assessed in pa allel
by exac ly iden ical p ocesses. He e we in oduce Vipie,
a web based i al di e si y popula ion ool accep ing as
inpu a se o iles om i ome me agenomics NGS
analyses o mul iple samples. He e we p esen he wo k-
low and esul s using NGS samples om Human
Mic obiome P ojec and o he me agenomics s udies.
Func ional on all mode n b owse s, he high pe o m-
ance pipeline is eely a ailable o academic usage.
Implemen a ion
Ou pipeline p ocesses de-mul iplexed pai ed FASTQ
iles, he mos ypical p oduc o me agenomics sequen-
cing. Se e al s eps a e hen pe o med in pa allel o all
samples: quali y con ol (QC), de-no o assembly o
pu a i e genomic con igs, axonomic classi ica ion o he
assembled con igs and o phan single on eads by pe -
o ming Blas que ies agains a local cus om i us da a-
base de i ed om Genbank, and inally emapping o
he sequencing eads on o e e ence sequences iden i ied
by his axonomic classi ica ion. De aul analysis pa ame-
e s can be easily modi ied (e.g. he QC s ingency, o
he de no o assembly algo i hm).
Depic ed in Fig. 1, Vipie pipeline uses mul i p ocesso
a chi ec u e wi h in eg a ion o Pos g eSQL o pe o m-
ance and da a managemen while p o iding secu ed
in e ac i e esul s and allowing web o m pa ame e s o
QC, assembly and sco ing. The indi idual pa ame e s
and i s de aul alues a e lis ed in he use guide. T im-
ming and quali y con ol a e pa ame e based applying
Galaxy p ojec u ili ies [13, 14]. We ha e in eg a ed lead-
ing de-no o assembly ools - Vel e [15], Me aVel e
[16], IDBA [17] and MEGAHIT (SOAPDENOVO) [18]
and ABySS [19]; hese me hods and ools a e u he
desc ibed and e iewed [5, 20–22]. Taxonomic iden i i-
ca ion is pe o med using BLAST [23] agains a local
NCBI da abase es ic ed o whole i us genomes. The
inal s ep o he pa allel analysis emaps he aw eads
using BWA [24] on o a lis o bes ma ches om he
BLAST que ies, and lis s he coun o o iginal eads
ma ching o each o hese e e ences. In cases whe e
eads ma ch equally well o mul iple i uses, he sco e is
di ided among such bes ma ches o exp ess impo an ly
he ambigui y in assigna ion o he mo i s sha ed among
i al axa, and he unce ain y o he p esen ly a ailable
classi ica ion.
De-no o con igs and eads ha do no ma ch o any
cu en ly known i us, op ionally il e ed o human
genome and known ibosomal DNA, can be e ie ed
o u he analysis as his ‘da k ma e ’o he i ome
p esumably con aining no el i uses. Ou pipeline allows
a di ec expo o hese unmapped eads owing o
h ee-s ep il e ing s a egy. Reads unma ched o known
i uses a e i s dep i ed o sequences ha ma ch o ibo-
somal DNA o bac e ial, a cheal and ungal o igin. This is
pe o med by emapping he eads by he BWA p og am
o da abases o 16S, 23S and 5S DNA (a copy o
p.ncbi.nlm.nih.go /genomes/TARGET, and a educed
da abase o 5S DNA h p://www.combio.pl/ na/) [25].
The nex s ep emaps he educed se o eads o he
human genome. This s ep yields he po en ial da k ma e
o he human genome, mixed wi h a small p opo ion o
bac e ial genomic DNA. Ou pipeline does no il e ou
hese bac e ial genomic eads, as hey may con ain no el
lysogenic (do man ) phages.
VIPIE’s e e ence i us da abase was buil om h ee
sou ces and clus e ing he sequences o he 97% le el o
iden i y u he educed he complexi y. Fi s , all i uses
we e downloaded om he e seq da abase a he NCBI
(h ps:// p.ncbi.nih.go / e seq/ elease/ i al/), and educed
o 97% iden i y by using he CD-HIT p og am (h ps://
gi hub.com/weizhongli/cdhi /[26]). Then, all i us se-
quences labeled as “comple e”,wi h he“ xid10239”
(supe kingdom Vi uses) in he “O gn” ield we e e ie ed
om Genbank. The que y e ie ed app oxima ely 80,000
sequences om he da abase, which we e subsequen ly
educed o he 97% simila i y by using he CD-HIT
p og am. Finally, simila ly o p e ious wo da abases,
phages we e me ged and clus e ed om he Eu opean
Bioin o ma ics Ins i u e (EBI) eposi o y ( p.ebi.ac.uk/
pub/da abases/ as a iles/embl_genomes/genomes/Phage/).
The web o m, in e ace dialogs and esul s a e
p og ammed o HTML5 s anda ds and using Ja aSc ip
and mode n, open sou ce Ja aSc ip lib a ies (h ps://
jque y.o g, h ps://da a ables.ne ) o b owse compa i-
bili y. Biopy hon [27] is used o sequencing pa sing and
o ma ing. Pa allel p ocessing is achie ed ia py hon
(h ps://www.py hon.o g) subp ocess module implemen-
a ion and uses Pos g eSQL (h ps://www.pos g esql.o g)
schema o job acking and esul s me ging. S anda d
SMTP lib a y is used o no i ica ion, hence he email egis-
a ion equi emen . Clus e ed hea maps a e implemen ed
Lin e al. BMC Genomics (2017) 18:378 Page 2 o 11
Table 1 Compa ison o he exis ing i ome pipelines ools
Pipeline Tool Vipie Vi omeScan [11] Vi usTAP [8] Vi ome [16] Me a i [14] Taxonome [6] Me aSho [12]
P ima y goal Pa allel analysis o
mul iple i al
me agenomes
om web and
sui ed o molecula
epidemiology s udies.
To p o ile i omes
using da abases o
exis ing euka yo ic
i uses wi hou
assembly.
Iden i ica ion o i uses
in a sample, a e a
ho ough elimina ion
o known non- i al
sequences.
Classi ica ion o all
pu a i e ORF ound
in a i al me agenome,
cha ac e iza ion o
i al communi ies.
Analysis o i ome,
di e si y me ics and
ma ke gene
phylogenies.
Ul a as me agenomics
analysis ocusing on
de ec ion o mic oo ganisms,
including i us and bac e ial.
Highly accu a e and
comp ehensi e wo k low
o hos -associa e
mic obiome classi ica ion
on mul iple samples.
Web based Yes. No. Yes. Yes (Flash equi ed). Yes. Yes. No.
Ou pu s In e ac i e able, plo s
and aw downloads.
Clus e ed hea maps
wi h dynamic g oup
assignmen e-plo s.
S a ic popula ion
pie cha s. Sample
based clus e ed
hea maps.
Con ig based hi s
and seamless web
BLAST in e ace.
Rich collec ion o
sample sou ce
i ome ORF and
sequence ca ego ies.
Compa a i e analysis
o i omes and
anno a ions including
ne wo ks, nonme ic
dis ance and ee maps.
In e ac i e pie cha s wi h
kingdoms in bins and also
imp essi e sunbu s la e
sub classi ie s.
A K ona g aph and
In e ac i e Taxonomy
HTML able along
wi h cs ile.
Sou ce da a Pai ed-end eads;
as q o ma .
Sinle-end o
pai ed-end eads;
as q o ma .
Pai ed-end eads.
Accep s also single-end
eads; as q o ma .
s , o as q; in ended
o he 454-gene a ed
me agenomes.
Reads (>300 bases)
o assembled con igs.
Pai ed-end eads in
as q and as a o ma s.
Pai ed-end eads in
as q o ma .
T imming
and il e ing
YES, as he i s s ep. YES, a e
selec ion o i al
eads, a he le el
o a bam ile.
YES, as he i s s ep. YES: quali y based;
duplica e il e ing;
con amina ion
No speci ied. No speci ied. YES, as he i s s ep.
De-no o
assembly
YES, a choice o
assemble s.
No. YES, a choice o
assemble s; done
a e sub ac ion s eps.
No. No. No. No.
Sub ac ion o
human e . and
bac e ial ibosomal
sequences
Op ional, only o
he ou pu o da k
ma e sequences.
YES, using Human
Bes Ma ch Tagge .
No o ibosomal.
YES, also o he hos
da abases a ailable
(mouse e c.).
No speci ied o
human. Ribosome
is emo ed using
BLAST agains
DNA db.
No speci ied. No sub ac ed bu
epo ed as pa
o de ec ion.
Yes, epo s iden i ica ion
o human hos eads
and bac e ial mappings.
Means o i us
iden i ica ion
(a) BLAST agains a
pan- i al da abase.
(b) Remapping o
o iginal eads o he
iden i ied candida es.
Mapping o he
membe s o he
i us da abase
using bow ie2 [24].
BLAST sea ch agains
he NCBI n da abase.
P o ein BLASTP upon
wo da abases. Se e al
ie s o classi ica ion
o he ORFs.
No speci ied. Taxonome Binne DB
wi h 21 bp kme s
unique iden i ie s o
known i uses.
Cus om simila i y
wo k low wi h
hamming dis ance.
Vi us da abase
o iden i ica ion
A cus om da abase
con aining 20759
human, animal, plan
and bac e ial i uses.
Euka yo ic i uses
only. Fou cus om
da abases a ailable
o download.
Speci ici y is
main ained by he
sub ac ion s eps
p io o assembly
and BLAST sea ch.
UniRe 100 pep ide
da abase, i e anno a ed
p o ein da abases,
Me aGenomes On-line.
GAAS ool (h ps://
sou ce o ge.ne /
p ojec s/gaas/).
Binne DB needs o
be buil using KAnalyze
[42](h ps://sou ce o ge.
ne /p ojec s/kanalyze/ iles/).
TANGO [43]and
NCBI Taxonomy [44].
Ac ion when a
ead maps o
di e en i uses
Sco e is spli among
he hi e e ence
sequences.
No speci ied. No speci ied. No speci ied. No speci ied. Assigns as ambiguous. Pa sed o human
endogenous e o i us
o he wise classi y as
ambiguous and
disca ded.
Mos ools use BLAST [23] o ini ial de ec ion o known e e ences. Vipie uniquely allows web pa allel analysis o mul i-samples and accoun s ead hi s o mul iple i al e e ences o comp ehensi e
popula ion p o iling
Lin e al. BMC Genomics (2017) 18:378 Page 3 o 11
wi h R ggplo 2 [28] while o he summa y and alpha di e -
si y s a is ics a e compu ed using cus om py hon sc ip s.
Popula ion maps and ead dis ibu ion coun summa y
cha s a e c ea ed using highcha s.js (h ps://www.high-
cha s.com) and cus om e en handle s o in e ac i i y.
Vipie is an ongoing open sou ced p ojec and a ailable a
h ps://sou ce o ge.ne /p ojec s/ ipie.
Resul s
Inpu samples and in e ac i e esul s
The pipeline u ili y is he e demons a ed on se o 11
samples whe e he inpu and esul s a e a ailable o all
use s. The sample se consis s o (a) blood, nasal, s ool
and agina da a om Human Me agenome P ojec
(HMP), (b) dia hea sample om gas oen e i is ou -
b eak (DRA004165 DNA Da a Bank Japan [29, 30]) used
in Vi usTAP and (c) s ool da a om in-house ongoing
A ican me agenomics p ojec [31, 32]. Table 2 lis s ele-
an accession iden i ie s, sou ces and numbe o eads
along wi h esul links. As he comp essed a chi ed
exceeds 1.2 gigaby es, a smalle subsampled a chi e
consis ing o 20% is a ailable o download on he home-
page and he o iginal comp essed FASTQ a chi ed is
a ailable on h ps://sou ce o ge.ne /p ojec s/ ipie/ iles/
da a [33]. End- o-end p ocessing o he 11 samples ook
82 min, p ocessing 29,778,980 eads ha includes
assembly, sco ing, and clus e ing and emo al o human
e e ence and known ibosomal e e ences. The pe -
o mance ime was measu ed a e he a chi e was
uploaded as ile upload depends ully on local ne wo k
speed. The in e ac i e esul s, wi h popula ion p o ile
maps and il e able i al hi ables a e accessible a :
h ps://bin .u a. i/ ipie/ esul s.h ml?key=eLZPuObVoU.
Resul links a e accessible wi hou egis a ion and
designed o be sha ed among collabo a o s whe eas job
his o y and ac i e jobs a e isible only o egis e ed in es-
iga o s. The esul s a e di ided in o panels o Popula ion
p o ile & g oup assignmen , QC & Da k ma e epo ,
Summa y & alpha di e si y, and Vi al hi s able. Raw
esul s, including unmapped da k ma e eads ha o no
ma ch o any known i us can be also downloaded.
Figu e 2 shows g oup-based popula ion pie cha s and
alpha di e si y as measu ed by Shannon en opy [34].
The popula ion pie cha sizes a e ela i e o o al num-
be o hi s and hei slices a e ully in e ac i e as clicking
on he slices a e ses he axonomy le els. The ool
ound 167 unique accessions ac oss he samples and an
easy o use sea chable and so able sample hi s able is
p o ided and bes expe ienced om he b owse , whe e
he able can be collapsed based on axonomy and
sample i al hi s can be downloaded as a ex ile eady
o Excel impo .
Ou use guide p o ides sc eensho s and di ec ions on
il e ing he sample hi s able and using he il e ing
unc ion, we ound Human He pes hi s on a HMP blood
sample SRS072276, whe e he pes in hema ological
samples ha e been epo ed in a p io mic obiome
and hema opoiesis epo [35]. Ou esul s showed
ha i us popula ion p o iles a e unique ac oss body
si es, epo ed also in Vi omeScan and isually shown
Fig. 1 Vipie web low cha . Fo e iciency, sample based pai ed FASTQ iles a e uploaded as a zipped a chi e wi h op ional mapping ile. Illumina
BaseSpace a chi e downloads can be used wi hou changes. All pipeline pa ame e s can be en e ed using he web o m. The de aul alues and
use case a e lis ed in he use guide a ailable a home page along wi h example mul i-sample a chi e inpu
Lin e al. BMC Genomics (2017) 18:378 Page 4 o 11
in he clus e ed maps. In e es ingly, in he s ool sam-
ple SRS012902, c Assphage [36] was by a he highes
i us de ec ed. Figu e 3 shows he clus e ed hea map gen-
e a ed in R, and i co ec ly clus e ed heal hy HMP sam-
ple ypes oge he [11] while Japanese gas oen e i is and
A ican samples showed p o oundly di e en signa u es.
Compa isons
We i s compa ed ou pe o mance o ha o Vi omeS-
can. While Vi omeScan s a es ha i suppo s mul iple
samples, i equi es local ins alla ion wi h 50+ gigaby es
o da abase equi emen s. The 20 HMP samples used o
i s alida ion, only he s ool samples passed QC [37] and
likely due o iming, he o he sample ypes we e no
a ailable on HMP download page. Ou summa y and
clus e indings o s ool samples and e oau icula , wi h
he highes di e si y, samples ag ee wi h Vi omeScan
and o he HMP indings o ~5.5 gene a pe sample [38].
We we e unable o ep oduce he he pes associa ions
epo ed wi h agina samples as hose samples a e no
longe a ailable. Inpu pa ame e s, in e ac i e maps, QC
epo (Fig. 4a) and i al hi s o he 11 samples a e
accessible a h ps://bin .u a. i/ ipie/ esul s.h ml?key=eLZ
PuObVoU and Table 2 con ains accession ids along wi h
sample ead sizes.
Then pe o mance o Vipie was compa ed o Vi usTAP.
I s web based de no o assembly dedica ed pipeline
equi ed 17 min o p ocess he DRA004165 sample om
a s udy o gas oen e i is [29] in Japan. Vi usTAP capably
de ec ed 11 Human o a i uses whe e his esul is ci ed
and also a ailable as i s example esul s. Vipie using he
same inpu de ec ed simila indings o 14 Human o a i-
uses s ains (shown in Addi ional ile 1: Use guide Figu e
10B) and also in e es ingly S ep ococcus phage s ains.
Using he same sample, ou pipeline equi ed 32 min due
o pos assembly emapping wi h cus om sco ing and hen
unmapped o igin il e ing. Because o Vipie’s pa allel com-
pu ing design, he a chi e o 11 samples and mo e han 10
imes he amoun o eads, ook jus 82 min. The mo e
comp ehensi e indings also highligh he sco ing spli
s a egy on ead hi s on mul iple i uses and in es iga ion
o unmapped i al ead o igins shown in Fig. 4b.
Fu he mo e, benchma king was assessed and com-
pa ed wi h he ecen ly published Me aSho , using i s
simula ed a i icial da ase wi h a e y high sha e o
human sequences mixed wi h low amoun s o many
di e en i al sequences. Table 3 below shows he simila
p ecision and ecall esul s o he wo ools. Vipie has a
sligh ly highe pe cen age o unclassi ied i al eads likely
due o subsampling o he ini ial da ase , and due o he
ac ha we op imized he i us BLAST da abase by e-
mo ing sequences ha we e less dis an han 3% om i s
closes ela i e; simila educ ion o axonomic complexi y
is known om e.g. bac e iome p o iling. The sc ip and
Vipie esul s used o compu ing his s a is ics a e a ail-
able wi h README in Vipie p ojec page on Sou ceFo ge.
We a e g a e ul o Me aSho au ho s o pe mission o
use hei simula ed da a, cons uc ed using ART [39].
Table 2 NGS samples used in Vipie alida ion om Human Mic obiome P ojec , A ica s udy, and dia hea sample sou ced in Japan
gas oen e i is ou b eak. Vi omeScan lis ed 20 HMP samples bu only S ool ypes o 4 samples passed QC
AccessionId Sou ce Sample Type Numbe o Reads
a
Sample used in Vipie-Vi omeScan-Vi usTAP alida ion Vipie Resul s
b
SRS072276 HMP Blood 438,879 Yes-No-No 1,2
SRS072318 HMP Blood 753,994 Yes-No-No 1,2
SRS019033 HMP Re oau icula 1,285,003 Yes-No-No 1
SRS016944 HMP Re oau icula 1,619,439 Yes-No-No 1
SRS012902 HMP S ool 2,039,473 Yes-Yes-No 1
SRS014923 HMP S ool 2,009,179 Yes-Yes-No 1
SRS014466 HMP Vagina 367,077 Yes-No-No 1,2
SRS015072 HMP Vagina 495,256 Yes-No-No 1,2
SRS072313 HMP Nasal 320,672 Yes-No-No 2
SRS072261 HMP Nasal 367,384 Yes-No-No 2
SRS072366 HMP Nasal 114,414 Yes-No-No 2
S11 A ica S ool 1,634,821 Yes-No-No 2
S12 A ica S ool 1,191,427 Yes-No-No 2
S14 A ica S ool 1,143,784 Yes-No-No 2
DRA004165 Japan Dia heal 1,108,688 Yes-No-Yes 2
In addi ion o hose s ool samples, Vipie es a chi e includes 4 o he HMP sample ypes. Resul links wi h pe o mance ime a e also p o ided
a
Inpu a chi e o Resul 2 samples (subsampled 20% 225 MB) a ailable a : h ps://bin .u a. i/ ipie/da a/ ipie_a chi e_ssampled.zip
b
Resul s 1: h ps://bin .u a. i/ ipie/ esul s.h ml?key=2HSPXukkDS (66 min)
Resul s 2: h ps://bin .u a. i/ ipie/ esul s.h ml?key=eLZPuObVoU (82 min)
Lin e al. BMC Genomics (2017) 18:378 Page 5 o 11

Discussion
Vipie in e ace is implemen ed wi h HTML5 s anda ds
and u ilizes open sou ce Ja aSc ip lib a ies. Unlike olde
and Adobe Flash based applica ions, Vipie does no
equi e addi ional ins alla ions and suppo s all mode n
HTML5 complian b owse s while o e ing a consis en
use expe ience. The inpu pa ame e o m is designed
o be clean and o g oup in o p ocessed componen s
whe e each elemen has cus om alida ion ules. The
componen de ails and ules a e lis ed in he use guide.
Secu ed and in e ac i e analysis esul s a e accessed wi h
enc yp ed links and o p omo e collabo a ion, can be
sha ed wi hou egis a ion. Sample based alpha di e si y
is p o ided, using Shannon en opy index [34] (Fig. 2) as
a ep esen a i e o di e si y me hods [35]. Vipie in ui-
i ely o e s web based, o m o ile upload sample g oup
Shannon en o
py
index
Unique accesso ies (log)
Alpha di e si y
Nasal
Vagina
A ica
Blood
Vi usTAP
-0.5 0 0.5 1 1.5 2 2.5 3 3.5
0
1
2
3
4
5
A
B
Fig. 2 In e ac i e popula ion p o ile maps and di e si y. Vipie esul s a e secu ely accessed and b owse based. aPopula ion cha slices a e
clickable and hei sizes ep esen ela i e pe cen age o ele an axonomy le el. Dia heal sample is domina ed by dsRNA (o ange) Ro a i us
while A ican s ool samples con ain ssRNA (g een) and dsDNA i uses. bAlpha di e si y is calcula ed using Shannon en opy. Vipie cha s a e
in e ac i e and can be sa ed as mul iple image o ma s
Lin e al. BMC Genomics (2017) 18:378 Page 6 o 11
S14
S11
SRS015072
SRS014466
SRS072261
SRS072366
SRS072313
S12
SRS072276
SRS072318
Dia hea
KF812551.1
NC_011222.1
NC_016770.1
KF726049.1
KJ870912.1
AY843304.1
KU355273.1
JQ173883.1
KP887098.1
KF726046.1
KF371693.1
KJ870927.1
DQ005111.1
KJ870919.1
KF726047.1
KR093640.1
KP343683.1
KF726054.1
KF371836.1
JX169867.1
JX169866.1
FJ647224.1
AC_000192.1
FJ647223.1
JF792616.1
KP266574.1
AY302543.1
AY302554.1
AY843298.1
NC_009514.1
AF324493.2
JX904130.1
KF726053.1
NC_019782.1
NC_011801.1
NC_008168.1
KT336321.1
NC_023503.1
KT336320.1
KP343840.1
NC_014094.1
KR093631.1
NC_025726.1
NC_019710.1
CP000711.1
AY208746.1
NC_022518.1
KJ716849.1
NC_014080.1
NC_027398.1
NC_018285.1
NC_005344.1
DQ902712.1
KP343864.1
KP343854.1
NC_002730.1
KP343844.1
NC_023984.1
JQ347801.1
KP343843.1
JN980171.1
AY184221.1
NC_014075.1
EU078592.1
KP343824.1
NC_014089.1
KP869108.1
NC_019716.1
HQ188292.1
NC_009225.1
AY302547.1
KP289439.1
DQ246620.1
EF174468.1
KF878966.1
KP290111.1
AY302556.1
KP289437.1
JX976771.1
KF042343.1
FJ357838.1
AY302544.1
EF174469.1
KC897073.1
JQ041368.1
AY302546.1
KP266575.1
KT353721.1
JX898907.1
JX476161.2
AY556070.1
KX810066.1
KR107057.1
EF634316.1
DQ534205.1
AY302542.1
AY673831.1
AY302553.1
JN596587.1
JQ729993.1
KF781525.1
AY843300.1
AY843306.1
JX476169.2
AY302548.1
JX476166.2
AF268065.1
Sample Vi al Hi s
310 1 2 3
Row Z Sco e
Colo Key
Nasal
Vagina
A ica
Blood
Vi usTAP
Fig. 3 (See legend on nex page.)
Lin e al. BMC Genomics (2017) 18:378 Page 7 o 11
(See igu e on p e ious page.)
Fig. 3 Clus e ed hea map o HMP, A ican and Japanese dia heal samples. Public NGS da a om di e en conso iums p o ide oppo uni ies o
ad anced compa a i e i ome analysis. Heal hy HMP sample ypes clus e ed co ec ly (nasal, aginal, blood samples) while a Japanese sample
(gas oen e i is da ase om he Vi usTAP epo ) and A ican samples (known o be posi i e o mul iple i uses) showed di e en signa u es.
HMP samples can be iden i ied using he legend on uppe igh , wi h oli e g een o nasal, yellow o agina and blue o blood. Samples om
u al A ica and Vi usTAP (Japan) a e ma ked in colo s b ick and ed
A
B
Fig. 4 QC and dis ibu ion o eads including da k i al ma e . aThe cha shows he numbe o NGS eads e ained pe sample h ough
QC, in e lacing and de no o assembly. bSample eads, along he x-axis and hei aligned o igins a e shown as s acked ba s. Shown in black,
unmapped i al ‘da k ma e ’is o high in e es ac oss i ology s udies. Blue ba s ep esen bac e ial ibosome, g een o human while ed is o
known i al ma ches
Lin e al. BMC Genomics (2017) 18:378 Page 8 o 11
eassignmen whe e popula ion and clus e ed maps a e
eanalyzed and dynamically ed awn. The pipeline p o-
duces a c oss abula ion simila o he ope a ional axo-
nomic uni (OTU) ables om bac e iome p o iling,
addi ional s a is ics is doable wi h ad ance R packages
such as phyloseq [40] and deseq2 [41].
O en, published pipelines emphasize ha hei pe -
o mance is by o de s o magni ude as e han exis ing
s a egies [7, 8] and ha he asks can be comple ed in
he o de o minu es o single hou s in a si ua ion whe e
exis ing i uses accoun only o a mino ac ion o he
o al ead coun . We belie e ha he p esen Vipie
pipeline o e s as da a p ocessing o mos ele an
applica ions, including eal- ime assessmen o i al
epe oi e in clinical samples. Fo compa ison, Vi usTAP
p ocessing, up o assembly wi h 1 sample (~2 million
eads, 172 MBs) ook 17 min (Inpu upload ime is no
included as i is dependen comple ely on local ne wo k
speed.). Vipie p ocess he same sample in 32 min includ-
ing assembly, c oss alida ion sco ing/ emapping, known
e e ence il e ing and i al da k ma e p ocessing.
Pa allel implemen a ion is ideal o mul i-sample p o-
cessing and inpu se o 11 samples (Table 2), consis ing
o ~30 million eads, 1.22 GBs comp essed and p oc-
essed in 82 min. The e is no concu en limi on he
numbe o samples eligible o p ocessing o he han a
small da abase o e head. Job comple ion ime has a di -
ec ela ionship o he sample wi h he highes ead
dep h and i is well known ha in e lacing and assembly
a e high memo y asks. The de no o assembly s ep im-
plemen s andom subsampling on use de ined ead pe -
cen age, de aul o 75% wi h a maximum o 1,000,000
NGS eads pe sample. Ve y la ge a chi es can su e
om ne wo k imeou s on ile upload. In o e coming
his scena io, we ha e success ully deployed Vipie on
clus e compu ing en i onmen and analyze housands
o samples consis ing o e aby es o da a using SLURM,
he de aul u ili y o Linux high pe o mance compu -
ing. We belie e ha ou s a egy o e s a good balance
be ween bea able algo i hm speed on mos machines,
and a ailabili y o mul iple sample p ocessing.
Impo an ly, he pipeline o e s a se o iles wi h bac-
e ial, human, and unknown sequences ( he “da k ma -
e ”o he i ome). Da k ma e eads a e he emaining
unmapped eads a e il e ing o human and bac e ial
ibosomes. I has been long known ha he unknown
da k ma e is ex emely aluable in i ome analysis [9]
and in ocus wi h he ecen disco e y o new bac e-
iophage i us c Assphage while i s bac e ial hos s ill
unknown [36]. Many componen s o his “da k ma e ”
o he i ome ha e been obse ed ac oss s udies, and a e
likely o ep esen exis ing i uses, ye hei axonomy is
p esen ly unknown. The lack o axonomic classi ica ion
howe e should no p eclude hei use as p o isional
en i ies, exposu es ha a e es able and quan i iable in
epidemiological s udies. Figu e 4b shows an in e ac i e
sample based cha consis ing o s acked ba s ep esen -
ing he pe cen age o eads mapped o human, bac e ial
ibosomes, known i uses and da k ma e . I is appa en
ha hese unmapped eads domina ed hese NGS sam-
ples and deepe ad anced analyses a e necessa y. As
such, i al da k ma e aw eads a e pa o downloads.
An o en-o e looked aspec is he unce ain y in i us
iden i ica ion. The Genbank da abase con ains many
simila isola es o almos e e y ele an i us se o ype.
This means ha mos eads o con igs would map o
mul iple di e en sequenced i us isola es. In single
sample s udies his does no pose any p oblem - he
axonomy is concluded as he highes sco ing hi , o he
i s o a se o simila ly high sco ing o ganisms. This
howe e canno be done when a pipeline p ocesses
mul iple samples a he same ime: due o he known
in insic a iabili y o he i uses, e en a single subjec
may p oduce wo di e en samples whe e di e en i us
quasi-species may p e ail ha will p e e en ially map o
wo di e en i us e e ence sequences. The e a e wo
possible solu ions o he p oblem: he Vi omeScan pipe-
line employed one whe e he da abases a e smalle wi h
a limi ed scope. Un o una ely, he s a egy owa ds hei
Table 3 (A) Read assignmen benchma k assessmen o
Me aSho and Vipie on simula ed da ase
a
consis ing o 19 582
500 human (94.5%), 986 114 bac e ial (4.8%) and 146 886 i al
(0.7%) eads. Vipie pe cen ages a e based on andom
subsampling o 1 000 000 eads and bac e ial s a is ics a e no
epo ed as Vipie epo s in o ma ion on bac e ial ibosome only
( he bac e ial genomic DNA is no il e ed ou , as i migh lead
o loss o do man phage sequences). (B) P ecision, Recall and
F-measu e a e calcula ed on he same da a. Inpu eads and
assessmen sc ip a e a ailable on Sou ceFo ge
b
A Assigned %
c
Co ec ly Assigned %
d
Me aSho Vipie Me aSho Vipie
Human (hos ) 99.18 99.27 99.99 99.27
Vi uses
Family 97.74 99.98 98.53 93.39
Genus 97.39 98.99 99.75 93.33
Species 97.81 93.66 96.70 92.97
B Human (hos ) Vi us
Me aSho Vipie Me aSho Vipie
P ecision (%) 100.00 100.00 98.30 96.85
Recall (%) 99.97 99.96 98.19 95.36
F-measu e (%) 100.00 99.98 98.07 96.08
Unclassi ied (%) 1.04 0.73 3.94 6.73
a
h ps:// ecascloud.ba.in n.i /index.php/s/nw4s9hqnF8QkBsK
b
h ps://sou ce o ge.ne /p ojec s/ ipie/ iles/ alida ion/k
c
The pe cen age e e s o he o al numbe o eads assignable o he speci ic
axonomic ank
d
The pe cen age e e s o he ele an assigned eads
Lin e al. BMC Genomics (2017) 18:378 Page 9 o 11