Genome size estimation from long read overlaps
Abstract
Poster for our work on genome size estimation from long read overlaps
Full text
Created with BioRender Poster Builder How big is your genome? Query Overlap Target 12 100 bp 100 bp possible overlaps actual overlaps 2 3Genome Size ≈ median ( ) Genome size estimation from long read overlaps Michael B Hall1, Chenxi Zhou2, Lachlan JM Coin1,3 1Department of Microbiology and Immunology, The University of Melbourne, at the Peter Doherty Institute for Infection and Immunity, 2Department of Genetics, The University of Cambridge, 3Department of Clinical Pathology, The University of Melbourne BACKGROUND ●Estimating genome size is important for understanding genome structure, determining sequencing coverage, and guiding downstream analyses such as assembly. ●Most tools use short-read data and rely on k-mer profiles, which don’t work well with longread technologies. ●We developed LRGE, a fast and accurate genome size estimator that works on long reads using read-to-read overlaps. METHODS Key points: ●Estimate genome size by analysing how many reads overlap each other. ●Reads with many overlaps suggest a smaller genome; fewer overlaps a larger genome. ●The median of these per-read estimates gives a genome size estimate with 92% CI. Tools & Implementation: ●Implemented in Rust, using minimap2 FFI bindings for efficient overlap detection. ●Two modes: ○2set: overlaps between small query and large target set. ○ava: all-v-all overlaps. RESULTS Absolute relative error (y-axis) for each method's (x-axis) genome size estimation on ONT and PacBio data. The yaxis is scaled according to a symmetric logarithm, which is linear between -1 and 1 and logarithmic (base 10) thereafter. The statistical annotations are the result of a Tukey's range test and are coloured by the sequencing platform being compared. Pairs with no annotation indicate no significant difference. The dashed lines in the violins are the quartiles. CPU time (left y-axis) and maximum memory usage (right y-axis) for each method (colours). The y-axis is logscaled (base 10). The statistical annotations are the result of a Tukey's range test and are coloured by the method with the lower mean value. Pairs with no annotation indicate no significant difference. The dashed lines in the violins are the quartiles. CONCLUSION ●LRGE is a lightweight, fast tool for genome size estimation from long read data. ●Outperforms k-mer-based tools, with lower computational cost than assembly. ●Useful for assembly and estimating coverage. Recommendation: Filter out lowquality and very short reads to improve estimates. Distribution of average read quality for samples with a relative error greater than 50% and all other samples. Impact of read length and quality on LRGE estimation accuracy. Lines are a linear regression fit to the points with the significance of the fit and the coefficient of determination (R2) in the legend. Infinite estimates are those reads which have no overlaps. ACKNOWLEDGEMENTS: This research was supported by the University of Melbourne’s Research Computing Services and the Petascale Campus Initiative. We thank Ryan Wick for insightful suggestions and Leah Roberts for discussions relating to large plasmid copy numbers. Poster