scieee AI-readable full text Open interactive document viewer

mzPeak presentation at HUPO2025

Klein, Joshua; Van Den Bossche, Tim; Wein, Samuel

Full text

mzPeak: a next generation MS run file format Background, challenges, questions for discussion Nov 11, 2025 – HUPO 2025, Toronto Joshua Klein, Tim Van Den Bossche, Samuel Wein What is a run format? File format to store spectra and (some) metadata acquired during an MS Experiment. Vendor implementations are what we traditionally call “raw files” Vendor formats are closed-source High density, fast load and store time (by first party software) Can change frequently with new tech and instruments First party software preferred/exclusive Open formats were developed to enable exchange of data outside of vendor ecosystem Allow comparison of data from different instruments Usable by any third party software Stable over long time span Converters to and from proprietary files exist (MSConvert, ThermoRawFileParser) Not always 1-to-1 mapping to vendor features and metadata Larger than vendor files Proprietary versus open formats For researchers −Interoperability between different platforms −Increases diversity of available software tools −Long-term data preservation For vendors −Decreases cost of internal development −Easier to bring new software products to market −Cloud and AI ready The importance of open formats Timeline of community effort Sep 2023 –Public discussion at HUPO (Busan, South Korea) Oct 2023 –Start MZNext mailing list Jan 2024 –Launch poll May 2024 –First internal version of technical whitepaper Jun 2024 –Public discussion at ASMS Feb 2025 –Broader whitepaper mzPeak uploaded to Zenodo, request comments on PREreview for public review Apr 2025 –discuss mzPeak project (white paper + implementation) at HUPO-PSI & First meeting of technical committee May 2025 –Submission of white paper to JPR October 2, 2025 – Publication of “mzPeak: Designing a Scalable, Interoperable, and FutureReady Mass Spectrometry Data Format” HUPO-PSI standard for storing experimental data Open XML-based standard First released in 2008, updated to 1.1.0 in 2009 mzML Size and read speed issues in mzML MS data files have grown several orders of magnitude in size (high res, ion mobility, etc) XML does not support random access (unless indexed) Low read and write speed Changes to the MS landscape Greater tolerance for non-”human-readable” formats Big tech led recognition of the importance of Open Source Scientist led recognition of the importance of FAIR science JPR article Submitted in May Published at the beginning of October Describes the goals of the project and a very general roadmap >2000 views so far Cloud nativeness Object storage is the future for large data mzPeak supports asynchronous IO over any backend that parquet supports Random access means that you only need to load the data you care about –saving hosting costs Open Questions What should be required to be included in an mzPeak file? What languages are a priority? Can we convince vendors to build their formats on top of mzPeak? How do we get software developers involved? Descriptive metadata for each spectrum spectra_metadata.mzpeak Descriptive metadata for each chromatogram or trace chromatograms_metadata.mzpeak Profile data defining the signal acquired for each chromatogram chromatograms_data.mzpeak Optionally, store centroids for any profile spectrum separately, e.g. if a vendor RAW contains peaks and profiles spectra_peaks.mzpeak Store any other files you want alongside the other data files as long as they do not end in “.mzpeak” or = “mzpeak_index.json” Any other files you like Anatomy of an mzPeak archive A file system directory, uncompressed ZIP archive, or a web address on the cloud *.mzpeak Profile or centroid data defining the signal acquired for each spectrum spectra_data.mzpeak Describe the files in the archive mzpeak_index.json Why Parquet •Columnar blocked encoding and compression. •Embedded indices make range queries fast and combining filters makes it even faster. •Flexible schemas let you add more columns without impacting other readers or having to sacrifice data types. •Stability and support across many, many languages via the Apache Foundation. •Strong industry buy-in from cloud providers and analytics companies. https://parquet.apache.org/docs/file-format/ Columnar blocked compression •Every field/column is compressed in chunks, but the compression and encoding can be tailored to the data type or content. •If you don’t want to read a column, you can skip all costs to read its bytes, let alone decoding and decompressing it. •There are several built-in encodings that are transparent, so automatic indices see the raw data, not the encoded data. •Byte shuffling for floating point numbers •Run length encoding for repetitive values •Dictionary encoding for large common values (or repetitive ones) •Support for popular fast compression algorithms like Zstandard Indices and sparse reads •Parquet supports range indices (min, max) at the row group and data page level. If you can isolate your target entries to a subset of them, you can skip all the other work. •Eagerly load one column to evaluate your constraints and use that to decide how to sparsely read the rest of your data. •By combining filters, e.g. start time > 5 AND start time <10 AND MS level = 1 makes it even more efficient. Highly encoded & compressed data pages Independent row group spectra_data spectra_metadata Parallel indices A better “lossy” compression for sparse data Example: Faster XIC extraction Example: Excellent Compression CREATE TABLE spectrum ( 'index' bigint, id varchar, ms_level smallint, 'time' real, polarity smallint, mz_signal_continuity varchar, spectrum_type varchar, number_of_data_points int, parameters parameter[], data_procesing_ref int, number_of_auxiliary_arrays int, auxiliary_arrays auxiliary_array[], mz_delta_model numeric[], total_ion_current numeric, base_peak_mz numeric, base_peak_intensity numeric, lowest_observed_mz numeric, highest_observed_mz numeric, ... ); CREATE TABLE selected_ion ( spectrum_index bigint, precursor_index bigint, selected_ion_mz numeric, charge_state int, intensity numeric, parameters parameter[], ... ); CREATE TABLE scan ( spectrum_index bigint, scan_start_time numeric, preset_scan_configuration int, filter_string varchar, ion_injection_time numeric, instrument_configuration_ref int, parameters parameter[], scan_windows scan_window[], ... ); CREATE TYPE param_value AS ( 'integer' integer, 'float' numeric, 'bool' boolean, 'string' varchar ); CREATE TYPE parameter AS ( 'name' varchar, accession varchar, 'value' param_value, unit varchar, ); Define multiple tables within metadata files. A minimal core of columns are required, but additional columns can be added. Columns may be structural or map onto controlled vocabulary terms using an inflection rule: <CVID>_<ACCESSION>_<NAME> CREATE TABLE point ( spectrum_index bigint, mz numeric, intensity numeric, ... ); #spectrum_metadata #spectrum_data Flexible schemas with a common core