Skim-sequencing for genomic selection in wheat: a comparison of marker platforms
Crain, Jared; Crossa, Jose; DeHaan, Lee; Dreisigacker, Susanne; Poland, Jesse; Singh, Ravi Prakash; Vitale, Paolo
- Publisher
- Zenodo
- Language
- en
Abstract
The promise of genomics-assisted breeding relies on efficient, affordable, and abundant molecular markers. Leveraging modern sequencing technology, commercial laboratory products, and open-source software, we demonstrate how ultra-low whole-genome sequencing coverage (skim-seq, 0.05-0.10x) can be a viable marker platform. The direct generation of sequence data followed by imputation provides an opportunity to implement genomic selection while being robust to future genomic (changes in reference genome) and technological (improvement in sequencing capacity) changes. We genotyped 1,709 wheat lines with genotyping-by-sequencing (GBS), a mid-density DArTAG SNP panel, and skim-seq (0.07x). All skim-seq samples were used to identify loci variants using a reference genome without the aid of any high-coverage samples. STITCH software was used for imputation to obtain 121,437 markers. Comparing high-confidence STITCH imputed loci (approximately 65,000 of 14M imputed loci) to high coverage samples resulted in the correct imputation for more than 97.5% of the markers. Using phenotypic data, a 5-fold cross validation was implemented for each marker platform. No one marker system performed the best in all test cases, with GBS often resulting in the highest correlation between observed and predicted values. The skim-seq correlations were typically within 0.03 of GBS, suggesting skim-seq can be a viable marker strategy for genomic prediction. As technology and computational pipelines advances, skim-seq appears to be a promising method to bridge the gap between targeted genotyping and whole-genome sequencing. The skim-seq method is highly flexible and can be optimized to a variety of program needs, potentially allowing for wide adoption by the plant breeding community.
Full text
Crain, J., Crossa, J., DeHaan, L., Dreisigacker, S., Poland, J., Singh, R.P., and Vitale, P. (2025). Skim-sequencing for genomic selection in wheat: a comparison of marker platforms. . Data Set: R scripts (as RMarkdown) for genomic analysis and development of skim-seq pipeline in wheat breeding with a comparison of three marker platforms. All RCode is documented in RMarkdown. The entire workflow progress from 0_0 to 4_0 sequentially, with the RMarkdown files documenting analysis and files. Original sequence fastq files are a part of the NCBI sequence read archive (SRA) (https://www.ncbi.nlm.nih.gov/bioproject/) BioProjects PRJNA498085 and part of umbrella project PRJNA901997, PRJNA1139475, and PRJNA1263701. Sequence files are only needed if single nucleotide polymorphisms (SNPs) are called using the GBS or skim-seq pipeline. All data for details in manuscript are included without needing original sequence data. Authors: Jared Crain Department of Plant Pathology, Kansas State University 4024 Throckmorton Plant Sciences Center, Manhattan, KS 66506, USA [email protected] Jared Crain (0000-0001-9484-8325) José Crossa (0000-0001-9429-5855) Lee DeHaan (0000-0002-6368-5241) Susanne Dreisigacker (0000-0002-3546-5989) Jesse Poland (0000-0002-7856-1399) Ravi P. Sign (0000-0002-4676-5071) Paolo Vitale (0000-0002-4353-5828) Files and Description: Files and folders located in the main directory and their contents: File or Directory Name Description of Contents README.pdf README file for workflow for manuscript tracking the skim-seq bioinformatics pipeline in wheat breeding.
Original_STITCH_Imputation_Unfiltered.tar Original STITCH imputation vcf files with samples, genotype and dosage calls. Separate .tar file Wheat_GS_SkimSeq.Rproj R Project for manuscript. 0_0_AWG_Wheat_SkimSeq_Demultiplex.Rmd Documents original Illumina demultiplexing 0_1_AWG_Wheat_Sample_Preparation_SkimSeq. Rmd Prepared fastq files for each sample for STITCH imputation. 0_2_AWG_Wheat_Variant_Discovery_SkimSeq.R md Identifies variants within 1,712 low coverage sequence wheat lines. 0_3_AWG_Wheat_GBS_SNP_Calling.Rmd Sets up GBS pipeline and filtering of GBS markers. 0_4_AWG_Wheat_DArTAG.Rmd Filters DArTAG markers and prepared them for genomic predciton. 0_5_AWG_Wheat_SNP_BCF_Unimputed.Rmd Evaluates low coverage sequencing including presence of reads and read depth in skim-seq. 1_0_AWG_Wheat_STITCH_Imputation.Rmd Documents STITCH imputation pipeline and bioinformatic filters for STITCH markers. 2_0_AWG_Wheat_High_Coverage_Samples.Rmd Evaluates STITCH imputed data to high coverage sequence calls for two samples. 3_0_AWG_Wheat_GS_CV.Rmd Implements genomic prediction using different marker data sets. 3_1_AWG_Wheat_GS_Comparison.Rmd Compiles genomic prediction for each marker platform. 3_2_AWG_Wheat_CV_Iteration.Rmd Sets up a 1,000 iterations of a random 5fold cross validation for all three marker platforms. 3_3_AWG_Wheat_Evaluation.Rmd Compiles results and evaluates standard error of the mean for data from 3_2_AWG_Wheat_CV_Iteration.Rmd. 3_4_AWG_Wheat_GWAS.Rmd Implements a genome-wide association study using phenotypic data with three different marker platforms. Compiles results and graphs output. 4_0_AWG_Wheat_Facts_Figures_Graphics.Rmd Develops information for manuscript and figures for manuscript. 0_0_AWG_Wheat_SkimSeq_Demultiplex.html Knitted Rmarkdown of each particular file. Provides package and software versions used in analysis. 0_1_AWG_Wheat_Sample_Preparation_SkimSeq. html 0_2_AWG_Wheat_Variant_Discovery_SkimSeq.ht ml 0_3_AWG_Wheat_GBS_SNP_Calling.html 0_4_AWG_Wheat_DArTAG.html
0_5_AWG_Wheat_SNP_BCF_Unimputed.html 1_0_AWG_Wheat_STITCH_Imputation.html 2_0_AWG_Wheat_High_Coverage_Samples.html 3_0_AWG_Wheat_GS_CV.html 3_1_AWG_Wheat_GS_Comparison.html 3_2_AWG_Wheat_CV_Iteration.html 3_3_AWG_Wheat_Evaluation.html 3_4_AWG_Wheat_GWAS.html 4_0_AWG_Wheat_Facts_Figures_Graphics.html File_List.txt Contains a list of files contained in the directory. For each file relative path and MD5 checksum is provided to verify against data corruption. MD5 calculated with md5 on Macintosh MacOS Sequoia Version 15.6.1. archive/ Files not used in final manuscript but may be of interest for applying skim-seq. Include a in silico pooling of samples for rapid variant identification. beocat/ Contain original genotyping-bysequencing (GBS) results, STITCH imputation, high coverage samples, and target file for STITCH imputation. data/ Directory that holds all original data, processed data, final files, and figures. scripts/ All scripts that were used to process files during pipeline processes or on HPC. out_error Bioinformatics output for sequence alignment, filtering, and quality scores. Note: All workflow can be traced using the Rmarkdown files in a sequential manner from 0_{sub_number}_Descriptor.Rmd to {max}_{sub_number}_File.Rmd. To run code, the working directory must be set to the main directory Wheat_GS_SkimSeq, which should set automatically if the Wheat_GS_SkimSeq.Rproj is opened in RStudio. Some files, particularly figures, that can be recreated easily from the code have been removed to reduce file size. To run shell scripts: shell scripts have been included, but file paths and directories must be updated for appropriate systems.