scieee Open visual document viewer

Vipie: web pipeline for parallel characterization of viral populations from multiple NGS samples

Lin, Jake,Kramka, Lenka,Autio, Reija,Hyöty, Heikki,Nykter, Matti,Cinek, Ondrej

Abstract

BioMed Central open access

Full text

SOFTWARE Open Access Vipie: web pipeline o pa allel cha ac e iza ion o i al popula ions om mul iple NGS samples Jake Lin 1 , Lenka K amna 2 , Reija Au io 3 , Heikki Hyö y 1,4* , Ma i Nyk e 1* and Ond ej Cinek 2* Abs ac Backg ound: Nex gene a ion sequencing (NGS) echnology allows labo a o ies o in es iga e i ome composi ion in clinical and en i onmen al samples in a cul u e-independen way. The e is a need o bioin o ma ic ools capable o pa allel p ocessing o i ome sequencing da a by exac ly iden ical me hods: his is especially impo an in s udies o mul i ac o ial diseases, o in pa allel compa ison o labo a o y p o ocols. Resul s: We ha e de eloped a web-based applica ion allowing di ec upload o sequences om mul iple i ome samples using cus om pa ame e s. The samples a e hen p ocessed in pa allel using an iden ical p o ocol, and can be easily eanalyzed. The pipeline pe o ms de-no o assembly, axonomic classi ica ion o i uses as well as sample analyses based on use -de ined g ouping ca ego ies. Tables o i us abundance a e p oduced om c oss- alida ion by emapping he sequencing eads o a union o all obse ed e e ence i uses. In addi ion, ead se s and epo s a e c ea ed a e p ocessing unmapped eads agains known human and bac e ial ibosome e e ences. Secu ed in e ac i e esul s a e dynamically plo ed wi h popula ion and di e si y cha s, clus e ed hea maps and a so able and sea chable abundance able. Conclusions: The Vipie web applica ion is a unique ool o mul i-sample me agenomic analysis o i al da a, p oducing sea chable hi s ables, in e ac i e popula ion maps, alpha di e si y measu es and clus e ed hea maps ha a e g ouped in applicable cus om sample ca ego ies. Known e e ences such as human genome and bac e ial ibosomal genes a e op ionally emo ed om unmapped (‘da k ma e ’) eads. Secu ed esul s a e accessible and sha eable on mode n b owse s. Vipie is a eely a ailable web-based ool whose code is open sou ce. Keywo ds: Me agenomics, Vi omes, Vi us, Assembly, NGS analysis, Visualiza ion, Pa allel p ocessing, Vi al da k ma e Backg ound The use o i ome me agenomics has been g owing apidly due o he inc easing demands o s udy he whole i ome in clinical samples and o e alua e he e olu ion o i al quasispecies du ing acu e and ch onic in ec ions. The applica ion o i ome sequencing ech- niques become use ul no only in in ec ious disease esea ch, bu also in associa ion s udies o p ima ily non-in ec ious condi ions, i.e. in diseases whe e he agen is p esumed o modi y he isk o he disease, which e ec is de ec able upon in es iga ion o a la ge numbe o subjec s only. These applica ions equi e an app oxima ion o i us quan i y, simila o wha has long been u ilized in bac e iome p o iling. As i uses lack a common sequence signa u e, me age- nomics sequencing o andom i al lib a ies emains he only easible way o an unbiased assessmen o he whole i ome. P esen ly, he need o accu a e quan i ica ion and in e p e a ion o i al popula ion me ics ac oss a se o samples c ea es a subs an ial challenge o his kind o me agenomics s udies. P ime obs acles o i ome in es iga o s a e he la ge gene ic he e ogenei y and also ha he majo i y o bioin o ma ic ools a e command line based and o e ly echnical, being com- pu a ionally demanding, wi h complica ed dependencies, * Co espondence: [email p o ec ed];ma i.nyk e[email p o ec ed]; [email p o ec ed] 1 BioMediTech and Facul y o Medicine and Li e Sciences, Uni e si y o Tampe e, PB 100FI-33014 Tampe e, Finland 2 Depa men o Pedia ics, 2nd Facul y o Medicine, Cha les Uni e si y and Uni e si y Hospi al Mo ol, V Ú alu 84, 150 06 P aha 5, Czech Republic Full lis o au ho in o ma ion is a ailable a he end o he a icle © The Au ho (s). 2017 Open Access This a icle is dis ibu ed unde he e ms o he C ea i e Commons A ibu ion 4.0 In e na ional License (h p://c ea i ecommons.o g/licenses/by/4.0/), which pe mi s un es ic ed use, dis ibu ion, and ep oduc ion in any medium, p o ided you gi e app op ia e c edi o he o iginal au ho (s) and he sou ce, p o ide a link o he C ea i e Commons license, and indica e i changes we e made. The C ea i e Commons Public Domain Dedica ion wai e (h p://c ea i ecommons.o g/publicdomain/ze o/1.0/) applies o he da a made a ailable in his a icle, unless o he wise s a ed. Lin e al. BMC Genomics (2017) 18:378 DOI 10.1186/s12864-017-3721-7 and p oducing ex based ou pu s ha a e no easily in e p e able [1–5]. Recen ly eleased web based applica- ions Taxonome [6], Vi usTAP [7], Vi ome [8] and Me a i [9, 10] ha e add essed some o he issues (espe- cially hose o use in e ac ion), bu mos ly ope a e only on single sample expe imen s wi h di e en wo k lows. Requi ing local dependencies and ins alla ion, Vi omeScan [11] and Me aSho [12] wo ks on mul iple samples. Some o hese ools we e designed o long (>300) eads o assembled con igs [8–10], which is limi ing as mode n me agenomics p ojec s including Human Mic obiome P ojec (HMP) [1, 2] p oduce mos ly high- h oughpu sho pai ed eads. Table 1 p o ides an o e iew o he p ima y ea u es and s a egies o hese di e en ools, including ou wo k. We aimed o open he possibili y o c ea ing a able o i al quan i ies o mul iple samples assessed in pa allel by exac ly iden ical p ocesses. He e we in oduce Vipie, a web based i al di e si y popula ion ool accep ing as inpu a se o iles om i ome me agenomics NGS analyses o mul iple samples. He e we p esen he wo k- low and esul s using NGS samples om Human Mic obiome P ojec and o he me agenomics s udies. Func ional on all mode n b owse s, he high pe o m- ance pipeline is eely a ailable o academic usage. Implemen a ion Ou pipeline p ocesses de-mul iplexed pai ed FASTQ iles, he mos ypical p oduc o me agenomics sequen- cing. Se e al s eps a e hen pe o med in pa allel o all samples: quali y con ol (QC), de-no o assembly o pu a i e genomic con igs, axonomic classi ica ion o he assembled con igs and o phan single on eads by pe - o ming Blas que ies agains a local cus om i us da a- base de i ed om Genbank, and inally emapping o he sequencing eads on o e e ence sequences iden i ied by his axonomic classi ica ion. De aul analysis pa ame- e s can be easily modi ied (e.g. he QC s ingency, o he de no o assembly algo i hm). Depic ed in Fig. 1, Vipie pipeline uses mul i p ocesso a chi ec u e wi h in eg a ion o Pos g eSQL o pe o m- ance and da a managemen while p o iding secu ed in e ac i e esul s and allowing web o m pa ame e s o QC, assembly and sco ing. The indi idual pa ame e s and i s de aul alues a e lis ed in he use guide. T im- ming and quali y con ol a e pa ame e based applying Galaxy p ojec u ili ies [13, 14]. We ha e in eg a ed lead- ing de-no o assembly ools - Vel e [15], Me aVel e [16], IDBA [17] and MEGAHIT (SOAPDENOVO) [18] and ABySS [19]; hese me hods and ools a e u he desc ibed and e iewed [5, 20–22]. Taxonomic iden i i- ca ion is pe o med using BLAST [23] agains a local NCBI da abase es ic ed o whole i us genomes. The inal s ep o he pa allel analysis emaps he aw eads using BWA [24] on o a lis o bes ma ches om he BLAST que ies, and lis s he coun o o iginal eads ma ching o each o hese e e ences. In cases whe e eads ma ch equally well o mul iple i uses, he sco e is di ided among such bes ma ches o exp ess impo an ly he ambigui y in assigna ion o he mo i s sha ed among i al axa, and he unce ain y o he p esen ly a ailable classi ica ion. De-no o con igs and eads ha do no ma ch o any cu en ly known i us, op ionally il e ed o human genome and known ibosomal DNA, can be e ie ed o u he analysis as his ‘da k ma e ’o he i ome p esumably con aining no el i uses. Ou pipeline allows a di ec expo o hese unmapped eads owing o h ee-s ep il e ing s a egy. Reads unma ched o known i uses a e i s dep i ed o sequences ha ma ch o ibo- somal DNA o bac e ial, a cheal and ungal o igin. This is pe o med by emapping he eads by he BWA p og am o da abases o 16S, 23S and 5S DNA (a copy o p.ncbi.nlm.nih.go /genomes/TARGET, and a educed da abase o 5S DNA h p://www.combio.pl/ na/) [25]. The nex s ep emaps he educed se o eads o he human genome. This s ep yields he po en ial da k ma e o he human genome, mixed wi h a small p opo ion o bac e ial genomic DNA. Ou pipeline does no il e ou hese bac e ial genomic eads, as hey may con ain no el lysogenic (do man ) phages. VIPIE’s e e ence i us da abase was buil om h ee sou ces and clus e ing he sequences o he 97% le el o iden i y u he educed he complexi y. Fi s , all i uses we e downloaded om he e seq da abase a he NCBI (h ps:// p.ncbi.nih.go / e seq/ elease/ i al/), and educed o 97% iden i y by using he CD-HIT p og am (h ps:// gi hub.com/weizhongli/cdhi /[26]). Then, all i us se- quences labeled as “comple e”,wi h he“ xid10239” (supe kingdom Vi uses) in he “O gn” ield we e e ie ed om Genbank. The que y e ie ed app oxima ely 80,000 sequences om he da abase, which we e subsequen ly educed o he 97% simila i y by using he CD-HIT p og am. Finally, simila ly o p e ious wo da abases, phages we e me ged and clus e ed om he Eu opean Bioin o ma ics Ins i u e (EBI) eposi o y ( p.ebi.ac.uk/ pub/da abases/ as a iles/embl_genomes/genomes/Phage/). The web o m, in e ace dialogs and esul s a e p og ammed o HTML5 s anda ds and using Ja aSc ip and mode n, open sou ce Ja aSc ip lib a ies (h ps:// jque y.o g, h ps://da a ables.ne ) o b owse compa i- bili y. Biopy hon [27] is used o sequencing pa sing and o ma ing. Pa allel p ocessing is achie ed ia py hon (h ps://www.py hon.o g) subp ocess module implemen- a ion and uses Pos g eSQL (h ps://www.pos g esql.o g) schema o job acking and esul s me ging. S anda d SMTP lib a y is used o no i ica ion, hence he email egis- a ion equi emen . Clus e ed hea maps a e implemen ed Lin e al. BMC Genomics (2017) 18:378 Page 2 o 11 Table 1 Compa ison o he exis ing i ome pipelines ools Pipeline Tool Vipie Vi omeScan [11] Vi usTAP [8] Vi ome [16] Me a i [14] Taxonome [6] Me aSho [12] P ima y goal Pa allel analysis o mul iple i al me agenomes om web and sui ed o molecula epidemiology s udies. To p o ile i omes using da abases o exis ing euka yo ic i uses wi hou assembly. Iden i ica ion o i uses in a sample, a e a ho ough elimina ion o known non- i al sequences. Classi ica ion o all pu a i e ORF ound in a i al me agenome, cha ac e iza ion o i al communi ies. Analysis o i ome, di e si y me ics and ma ke gene phylogenies. Ul a as me agenomics analysis ocusing on de ec ion o mic oo ganisms, including i us and bac e ial. Highly accu a e and comp ehensi e wo k low o hos -associa e mic obiome classi ica ion on mul iple samples. Web based Yes. No. Yes. Yes (Flash equi ed). Yes. Yes. No. Ou pu s In e ac i e able, plo s and aw downloads. Clus e ed hea maps wi h dynamic g oup assignmen e-plo s. S a ic popula ion pie cha s. Sample based clus e ed hea maps. Con ig based hi s and seamless web BLAST in e ace. Rich collec ion o sample sou ce i ome ORF and sequence ca ego ies. Compa a i e analysis o i omes and anno a ions including ne wo ks, nonme ic dis ance and ee maps. In e ac i e pie cha s wi h kingdoms in bins and also imp essi e sunbu s la e sub classi ie s. A K ona g aph and In e ac i e Taxonomy HTML able along wi h cs ile. Sou ce da a Pai ed-end eads; as q o ma . Sinle-end o pai ed-end eads; as q o ma . Pai ed-end eads. Accep s also single-end eads; as q o ma . s , o as q; in ended o he 454-gene a ed me agenomes. Reads (>300 bases) o assembled con igs. Pai ed-end eads in as q and as a o ma s. Pai ed-end eads in as q o ma . T imming and il e ing YES, as he i s s ep. YES, a e selec ion o i al eads, a he le el o a bam ile. YES, as he i s s ep. YES: quali y based; duplica e il e ing; con amina ion No speci ied. No speci ied. YES, as he i s s ep. De-no o assembly YES, a choice o assemble s. No. YES, a choice o assemble s; done a e sub ac ion s eps. No. No. No. No. Sub ac ion o human e . and bac e ial ibosomal sequences Op ional, only o he ou pu o da k ma e sequences. YES, using Human Bes Ma ch Tagge . No o ibosomal. YES, also o he hos da abases a ailable (mouse e c.). No speci ied o human. Ribosome is emo ed using BLAST agains DNA db. No speci ied. No sub ac ed bu epo ed as pa o de ec ion. Yes, epo s iden i ica ion o human hos eads and bac e ial mappings. Means o i us iden i ica ion (a) BLAST agains a pan- i al da abase. (b) Remapping o o iginal eads o he iden i ied candida es. Mapping o he membe s o he i us da abase using bow ie2 [24]. BLAST sea ch agains he NCBI n da abase. P o ein BLASTP upon wo da abases. Se e al ie s o classi ica ion o he ORFs. No speci ied. Taxonome Binne DB wi h 21 bp kme s unique iden i ie s o known i uses. Cus om simila i y wo k low wi h hamming dis ance. Vi us da abase o iden i ica ion A cus om da abase con aining 20759 human, animal, plan and bac e ial i uses. Euka yo ic i uses only. Fou cus om da abases a ailable o download. Speci ici y is main ained by he sub ac ion s eps p io o assembly and BLAST sea ch. UniRe 100 pep ide da abase, i e anno a ed p o ein da abases, Me aGenomes On-line. GAAS ool (h ps:// sou ce o ge.ne / p ojec s/gaas/). Binne DB needs o be buil using KAnalyze [42](h ps://sou ce o ge. ne /p ojec s/kanalyze/ iles/). TANGO [43]and NCBI Taxonomy [44]. Ac ion when a ead maps o di e en i uses Sco e is spli among he hi e e ence sequences. No speci ied. No speci ied. No speci ied. No speci ied. Assigns as ambiguous. Pa sed o human endogenous e o i us o he wise classi y as ambiguous and disca ded. Mos ools use BLAST [23] o ini ial de ec ion o known e e ences. Vipie uniquely allows web pa allel analysis o mul i-samples and accoun s ead hi s o mul iple i al e e ences o comp ehensi e popula ion p o iling Lin e al. BMC Genomics (2017) 18:378 Page 3 o 11 wi h R ggplo 2 [28] while o he summa y and alpha di e - si y s a is ics a e compu ed using cus om py hon sc ip s. Popula ion maps and ead dis ibu ion coun summa y cha s a e c ea ed using highcha s.js (h ps://www.high- cha s.com) and cus om e en handle s o in e ac i i y. Vipie is an ongoing open sou ced p ojec and a ailable a h ps://sou ce o ge.ne /p ojec s/ ipie. Resul s Inpu samples and in e ac i e esul s The pipeline u ili y is he e demons a ed on se o 11 samples whe e he inpu and esul s a e a ailable o all use s. The sample se consis s o (a) blood, nasal, s ool and agina da a om Human Me agenome P ojec (HMP), (b) dia hea sample om gas oen e i is ou - b eak (DRA004165 DNA Da a Bank Japan [29, 30]) used in Vi usTAP and (c) s ool da a om in-house ongoing A ican me agenomics p ojec [31, 32]. Table 2 lis s ele- an accession iden i ie s, sou ces and numbe o eads along wi h esul links. As he comp essed a chi ed exceeds 1.2 gigaby es, a smalle subsampled a chi e consis ing o 20% is a ailable o download on he home- page and he o iginal comp essed FASTQ a chi ed is a ailable on h ps://sou ce o ge.ne /p ojec s/ ipie/ iles/ da a [33]. End- o-end p ocessing o he 11 samples ook 82 min, p ocessing 29,778,980 eads ha includes assembly, sco ing, and clus e ing and emo al o human e e ence and known ibosomal e e ences. The pe - o mance ime was measu ed a e he a chi e was uploaded as ile upload depends ully on local ne wo k speed. The in e ac i e esul s, wi h popula ion p o ile maps and il e able i al hi ables a e accessible a : h ps://bin .u a. i/ ipie/ esul s.h ml?key=eLZPuObVoU. Resul links a e accessible wi hou egis a ion and designed o be sha ed among collabo a o s whe eas job his o y and ac i e jobs a e isible only o egis e ed in es- iga o s. The esul s a e di ided in o panels o Popula ion p o ile & g oup assignmen , QC & Da k ma e epo , Summa y & alpha di e si y, and Vi al hi s able. Raw esul s, including unmapped da k ma e eads ha o no ma ch o any known i us can be also downloaded. Figu e 2 shows g oup-based popula ion pie cha s and alpha di e si y as measu ed by Shannon en opy [34]. The popula ion pie cha sizes a e ela i e o o al num- be o hi s and hei slices a e ully in e ac i e as clicking on he slices a e ses he axonomy le els. The ool ound 167 unique accessions ac oss he samples and an easy o use sea chable and so able sample hi s able is p o ided and bes expe ienced om he b owse , whe e he able can be collapsed based on axonomy and sample i al hi s can be downloaded as a ex ile eady o Excel impo . Ou use guide p o ides sc eensho s and di ec ions on il e ing he sample hi s able and using he il e ing unc ion, we ound Human He pes hi s on a HMP blood sample SRS072276, whe e he pes in hema ological samples ha e been epo ed in a p io mic obiome and hema opoiesis epo [35]. Ou esul s showed ha i us popula ion p o iles a e unique ac oss body si es, epo ed also in Vi omeScan and isually shown Fig. 1 Vipie web low cha . Fo e iciency, sample based pai ed FASTQ iles a e uploaded as a zipped a chi e wi h op ional mapping ile. Illumina BaseSpace a chi e downloads can be used wi hou changes. All pipeline pa ame e s can be en e ed using he web o m. The de aul alues and use case a e lis ed in he use guide a ailable a home page along wi h example mul i-sample a chi e inpu Lin e al. BMC Genomics (2017) 18:378 Page 4 o 11 in he clus e ed maps. In e es ingly, in he s ool sam- ple SRS012902, c Assphage [36] was by a he highes i us de ec ed. Figu e 3 shows he clus e ed hea map gen- e a ed in R, and i co ec ly clus e ed heal hy HMP sam- ple ypes oge he [11] while Japanese gas oen e i is and A ican samples showed p o oundly di e en signa u es. Compa isons We i s compa ed ou pe o mance o ha o Vi omeS- can. While Vi omeScan s a es ha i suppo s mul iple samples, i equi es local ins alla ion wi h 50+ gigaby es o da abase equi emen s. The 20 HMP samples used o i s alida ion, only he s ool samples passed QC [37] and likely due o iming, he o he sample ypes we e no a ailable on HMP download page. Ou summa y and clus e indings o s ool samples and e oau icula , wi h he highes di e si y, samples ag ee wi h Vi omeScan and o he HMP indings o ~5.5 gene a pe sample [38]. We we e unable o ep oduce he he pes associa ions epo ed wi h agina samples as hose samples a e no longe a ailable. Inpu pa ame e s, in e ac i e maps, QC epo (Fig. 4a) and i al hi s o he 11 samples a e accessible a h ps://bin .u a. i/ ipie/ esul s.h ml?key=eLZ PuObVoU and Table 2 con ains accession ids along wi h sample ead sizes. Then pe o mance o Vipie was compa ed o Vi usTAP. I s web based de no o assembly dedica ed pipeline equi ed 17 min o p ocess he DRA004165 sample om a s udy o gas oen e i is [29] in Japan. Vi usTAP capably de ec ed 11 Human o a i uses whe e his esul is ci ed and also a ailable as i s example esul s. Vipie using he same inpu de ec ed simila indings o 14 Human o a i- uses s ains (shown in Addi ional ile 1: Use guide Figu e 10B) and also in e es ingly S ep ococcus phage s ains. Using he same sample, ou pipeline equi ed 32 min due o pos assembly emapping wi h cus om sco ing and hen unmapped o igin il e ing. Because o Vipie’s pa allel com- pu ing design, he a chi e o 11 samples and mo e han 10 imes he amoun o eads, ook jus 82 min. The mo e comp ehensi e indings also highligh he sco ing spli s a egy on ead hi s on mul iple i uses and in es iga ion o unmapped i al ead o igins shown in Fig. 4b. Fu he mo e, benchma king was assessed and com- pa ed wi h he ecen ly published Me aSho , using i s simula ed a i icial da ase wi h a e y high sha e o human sequences mixed wi h low amoun s o many di e en i al sequences. Table 3 below shows he simila p ecision and ecall esul s o he wo ools. Vipie has a sligh ly highe pe cen age o unclassi ied i al eads likely due o subsampling o he ini ial da ase , and due o he ac ha we op imized he i us BLAST da abase by e- mo ing sequences ha we e less dis an han 3% om i s closes ela i e; simila educ ion o axonomic complexi y is known om e.g. bac e iome p o iling. The sc ip and Vipie esul s used o compu ing his s a is ics a e a ail- able wi h README in Vipie p ojec page on Sou ceFo ge. We a e g a e ul o Me aSho au ho s o pe mission o use hei simula ed da a, cons uc ed using ART [39]. Table 2 NGS samples used in Vipie alida ion om Human Mic obiome P ojec , A ica s udy, and dia hea sample sou ced in Japan gas oen e i is ou b eak. Vi omeScan lis ed 20 HMP samples bu only S ool ypes o 4 samples passed QC AccessionId Sou ce Sample Type Numbe o Reads a Sample used in Vipie-Vi omeScan-Vi usTAP alida ion Vipie Resul s b SRS072276 HMP Blood 438,879 Yes-No-No 1,2 SRS072318 HMP Blood 753,994 Yes-No-No 1,2 SRS019033 HMP Re oau icula 1,285,003 Yes-No-No 1 SRS016944 HMP Re oau icula 1,619,439 Yes-No-No 1 SRS012902 HMP S ool 2,039,473 Yes-Yes-No 1 SRS014923 HMP S ool 2,009,179 Yes-Yes-No 1 SRS014466 HMP Vagina 367,077 Yes-No-No 1,2 SRS015072 HMP Vagina 495,256 Yes-No-No 1,2 SRS072313 HMP Nasal 320,672 Yes-No-No 2 SRS072261 HMP Nasal 367,384 Yes-No-No 2 SRS072366 HMP Nasal 114,414 Yes-No-No 2 S11 A ica S ool 1,634,821 Yes-No-No 2 S12 A ica S ool 1,191,427 Yes-No-No 2 S14 A ica S ool 1,143,784 Yes-No-No 2 DRA004165 Japan Dia heal 1,108,688 Yes-No-Yes 2 In addi ion o hose s ool samples, Vipie es a chi e includes 4 o he HMP sample ypes. Resul links wi h pe o mance ime a e also p o ided a Inpu a chi e o Resul 2 samples (subsampled 20% 225 MB) a ailable a : h ps://bin .u a. i/ ipie/da a/ ipie_a chi e_ssampled.zip b Resul s 1: h ps://bin .u a. i/ ipie/ esul s.h ml?key=2HSPXukkDS (66 min) Resul s 2: h ps://bin .u a. i/ ipie/ esul s.h ml?key=eLZPuObVoU (82 min) Lin e al. BMC Genomics (2017) 18:378 Page 5 o 11 Discussion Vipie in e ace is implemen ed wi h HTML5 s anda ds and u ilizes open sou ce Ja aSc ip lib a ies. Unlike olde and Adobe Flash based applica ions, Vipie does no equi e addi ional ins alla ions and suppo s all mode n HTML5 complian b owse s while o e ing a consis en use expe ience. The inpu pa ame e o m is designed o be clean and o g oup in o p ocessed componen s whe e each elemen has cus om alida ion ules. The componen de ails and ules a e lis ed in he use guide. Secu ed and in e ac i e analysis esul s a e accessed wi h enc yp ed links and o p omo e collabo a ion, can be sha ed wi hou egis a ion. Sample based alpha di e si y is p o ided, using Shannon en opy index [34] (Fig. 2) as a ep esen a i e o di e si y me hods [35]. Vipie in ui- i ely o e s web based, o m o ile upload sample g oup Shannon en o py index Unique accesso ies (log) Alpha di e si y Nasal Vagina A ica Blood Vi usTAP -0.5 0 0.5 1 1.5 2 2.5 3 3.5 0 1 2 3 4 5 A B Fig. 2 In e ac i e popula ion p o ile maps and di e si y. Vipie esul s a e secu ely accessed and b owse based. aPopula ion cha slices a e clickable and hei sizes ep esen ela i e pe cen age o ele an axonomy le el. Dia heal sample is domina ed by dsRNA (o ange) Ro a i us while A ican s ool samples con ain ssRNA (g een) and dsDNA i uses. bAlpha di e si y is calcula ed using Shannon en opy. Vipie cha s a e in e ac i e and can be sa ed as mul iple image o ma s Lin e al. BMC Genomics (2017) 18:378 Page 6 o 11 S14 S11 SRS015072 SRS014466 SRS072261 SRS072366 SRS072313 S12 SRS072276 SRS072318 Dia hea KF812551.1 NC_011222.1 NC_016770.1 KF726049.1 KJ870912.1 AY843304.1 KU355273.1 JQ173883.1 KP887098.1 KF726046.1 KF371693.1 KJ870927.1 DQ005111.1 KJ870919.1 KF726047.1 KR093640.1 KP343683.1 KF726054.1 KF371836.1 JX169867.1 JX169866.1 FJ647224.1 AC_000192.1 FJ647223.1 JF792616.1 KP266574.1 AY302543.1 AY302554.1 AY843298.1 NC_009514.1 AF324493.2 JX904130.1 KF726053.1 NC_019782.1 NC_011801.1 NC_008168.1 KT336321.1 NC_023503.1 KT336320.1 KP343840.1 NC_014094.1 KR093631.1 NC_025726.1 NC_019710.1 CP000711.1 AY208746.1 NC_022518.1 KJ716849.1 NC_014080.1 NC_027398.1 NC_018285.1 NC_005344.1 DQ902712.1 KP343864.1 KP343854.1 NC_002730.1 KP343844.1 NC_023984.1 JQ347801.1 KP343843.1 JN980171.1 AY184221.1 NC_014075.1 EU078592.1 KP343824.1 NC_014089.1 KP869108.1 NC_019716.1 HQ188292.1 NC_009225.1 AY302547.1 KP289439.1 DQ246620.1 EF174468.1 KF878966.1 KP290111.1 AY302556.1 KP289437.1 JX976771.1 KF042343.1 FJ357838.1 AY302544.1 EF174469.1 KC897073.1 JQ041368.1 AY302546.1 KP266575.1 KT353721.1 JX898907.1 JX476161.2 AY556070.1 KX810066.1 KR107057.1 EF634316.1 DQ534205.1 AY302542.1 AY673831.1 AY302553.1 JN596587.1 JQ729993.1 KF781525.1 AY843300.1 AY843306.1 JX476169.2 AY302548.1 JX476166.2 AF268065.1 Sample Vi al Hi s 310 1 2 3 Row Z Sco e Colo Key Nasal Vagina A ica Blood Vi usTAP Fig. 3 (See legend on nex page.) Lin e al. BMC Genomics (2017) 18:378 Page 7 o 11 (See igu e on p e ious page.) Fig. 3 Clus e ed hea map o HMP, A ican and Japanese dia heal samples. Public NGS da a om di e en conso iums p o ide oppo uni ies o ad anced compa a i e i ome analysis. Heal hy HMP sample ypes clus e ed co ec ly (nasal, aginal, blood samples) while a Japanese sample (gas oen e i is da ase om he Vi usTAP epo ) and A ican samples (known o be posi i e o mul iple i uses) showed di e en signa u es. HMP samples can be iden i ied using he legend on uppe igh , wi h oli e g een o nasal, yellow o agina and blue o blood. Samples om u al A ica and Vi usTAP (Japan) a e ma ked in colo s b ick and ed A B Fig. 4 QC and dis ibu ion o eads including da k i al ma e . aThe cha shows he numbe o NGS eads e ained pe sample h ough QC, in e lacing and de no o assembly. bSample eads, along he x-axis and hei aligned o igins a e shown as s acked ba s. Shown in black, unmapped i al ‘da k ma e ’is o high in e es ac oss i ology s udies. Blue ba s ep esen bac e ial ibosome, g een o human while ed is o known i al ma ches Lin e al. BMC Genomics (2017) 18:378 Page 8 o 11 eassignmen whe e popula ion and clus e ed maps a e eanalyzed and dynamically ed awn. The pipeline p o- duces a c oss abula ion simila o he ope a ional axo- nomic uni (OTU) ables om bac e iome p o iling, addi ional s a is ics is doable wi h ad ance R packages such as phyloseq [40] and deseq2 [41]. O en, published pipelines emphasize ha hei pe - o mance is by o de s o magni ude as e han exis ing s a egies [7, 8] and ha he asks can be comple ed in he o de o minu es o single hou s in a si ua ion whe e exis ing i uses accoun only o a mino ac ion o he o al ead coun . We belie e ha he p esen Vipie pipeline o e s as da a p ocessing o mos ele an applica ions, including eal- ime assessmen o i al epe oi e in clinical samples. Fo compa ison, Vi usTAP p ocessing, up o assembly wi h 1 sample (~2 million eads, 172 MBs) ook 17 min (Inpu upload ime is no included as i is dependen comple ely on local ne wo k speed.). Vipie p ocess he same sample in 32 min includ- ing assembly, c oss alida ion sco ing/ emapping, known e e ence il e ing and i al da k ma e p ocessing. Pa allel implemen a ion is ideal o mul i-sample p o- cessing and inpu se o 11 samples (Table 2), consis ing o ~30 million eads, 1.22 GBs comp essed and p oc- essed in 82 min. The e is no concu en limi on he numbe o samples eligible o p ocessing o he han a small da abase o e head. Job comple ion ime has a di - ec ela ionship o he sample wi h he highes ead dep h and i is well known ha in e lacing and assembly a e high memo y asks. The de no o assembly s ep im- plemen s andom subsampling on use de ined ead pe - cen age, de aul o 75% wi h a maximum o 1,000,000 NGS eads pe sample. Ve y la ge a chi es can su e om ne wo k imeou s on ile upload. In o e coming his scena io, we ha e success ully deployed Vipie on clus e compu ing en i onmen and analyze housands o samples consis ing o e aby es o da a using SLURM, he de aul u ili y o Linux high pe o mance compu - ing. We belie e ha ou s a egy o e s a good balance be ween bea able algo i hm speed on mos machines, and a ailabili y o mul iple sample p ocessing. Impo an ly, he pipeline o e s a se o iles wi h bac- e ial, human, and unknown sequences ( he “da k ma - e ”o he i ome). Da k ma e eads a e he emaining unmapped eads a e il e ing o human and bac e ial ibosomes. I has been long known ha he unknown da k ma e is ex emely aluable in i ome analysis [9] and in ocus wi h he ecen disco e y o new bac e- iophage i us c Assphage while i s bac e ial hos s ill unknown [36]. Many componen s o his “da k ma e ” o he i ome ha e been obse ed ac oss s udies, and a e likely o ep esen exis ing i uses, ye hei axonomy is p esen ly unknown. The lack o axonomic classi ica ion howe e should no p eclude hei use as p o isional en i ies, exposu es ha a e es able and quan i iable in epidemiological s udies. Figu e 4b shows an in e ac i e sample based cha consis ing o s acked ba s ep esen - ing he pe cen age o eads mapped o human, bac e ial ibosomes, known i uses and da k ma e . I is appa en ha hese unmapped eads domina ed hese NGS sam- ples and deepe ad anced analyses a e necessa y. As such, i al da k ma e aw eads a e pa o downloads. An o en-o e looked aspec is he unce ain y in i us iden i ica ion. The Genbank da abase con ains many simila isola es o almos e e y ele an i us se o ype. This means ha mos eads o con igs would map o mul iple di e en sequenced i us isola es. In single sample s udies his does no pose any p oblem - he axonomy is concluded as he highes sco ing hi , o he i s o a se o simila ly high sco ing o ganisms. This howe e canno be done when a pipeline p ocesses mul iple samples a he same ime: due o he known in insic a iabili y o he i uses, e en a single subjec may p oduce wo di e en samples whe e di e en i us quasi-species may p e ail ha will p e e en ially map o wo di e en i us e e ence sequences. The e a e wo possible solu ions o he p oblem: he Vi omeScan pipe- line employed one whe e he da abases a e smalle wi h a limi ed scope. Un o una ely, he s a egy owa ds hei Table 3 (A) Read assignmen benchma k assessmen o Me aSho and Vipie on simula ed da ase a consis ing o 19 582 500 human (94.5%), 986 114 bac e ial (4.8%) and 146 886 i al (0.7%) eads. Vipie pe cen ages a e based on andom subsampling o 1 000 000 eads and bac e ial s a is ics a e no epo ed as Vipie epo s in o ma ion on bac e ial ibosome only ( he bac e ial genomic DNA is no il e ed ou , as i migh lead o loss o do man phage sequences). (B) P ecision, Recall and F-measu e a e calcula ed on he same da a. Inpu eads and assessmen sc ip a e a ailable on Sou ceFo ge b A Assigned % c Co ec ly Assigned % d Me aSho Vipie Me aSho Vipie Human (hos ) 99.18 99.27 99.99 99.27 Vi uses Family 97.74 99.98 98.53 93.39 Genus 97.39 98.99 99.75 93.33 Species 97.81 93.66 96.70 92.97 B Human (hos ) Vi us Me aSho Vipie Me aSho Vipie P ecision (%) 100.00 100.00 98.30 96.85 Recall (%) 99.97 99.96 98.19 95.36 F-measu e (%) 100.00 99.98 98.07 96.08 Unclassi ied (%) 1.04 0.73 3.94 6.73 a h ps:// ecascloud.ba.in n.i /index.php/s/nw4s9hqnF8QkBsK b h ps://sou ce o ge.ne /p ojec s/ ipie/ iles/ alida ion/k c The pe cen age e e s o he o al numbe o eads assignable o he speci ic axonomic ank d The pe cen age e e s o he ele an assigned eads Lin e al. BMC Genomics (2017) 18:378 Page 9 o 11