Modular comparison of untargeted metabolomics processing steps
Модульное сравнение этапов обработки в нетаргетном метаболомике
2024-11-27
SCID: 54.1/32wb2k2d
Discuss with AI
Compound DiscovererMS-DIALMZmineXCMSanion exchange chromatographyblank filteringdata scaling (auto scaling)feature detectionhigh resolution mass spectrometrymanual integrationmissing value imputationuntargeted metabolomics
Figures from the paper
Abstract (AI)
Untargeted metabolomics requires robust and reliable strategies for data processing to extract relevant information form the underlying raw data. Multiple platforms for data processing are available, but the choice of software tool can have an impact on the analysis. This study provides a comprehensive evaluation of four workflows based on commonly used metabolomics software tools: XCMS, Compound Discoverer, MS-DIAL, and MZmine. These tools were applied to a dataset derived from bovine saliva samples spiked with small polar molecules analyzed by anion exchange chromatography coupled to high resolution mass spectrometry. The analysis revealed significant differences in the number and overlap of detected features, with only approximately 8 % of the features included in all four peak tables. Among the overlapping features, MS-DIAL demonstrated the greatest similarity to manual integration, while XCMS and MZmine also performed well. In contrast, Compound Discoverer had issues to reliably integrate high baseline peaks. This study also explores various post-processing strategies, including missing value imputation, transformation, scaling, and filtering. The assessment of missing values indicated that they primarily originated from low abundance, making imputation with small values the most effective approach. No clear evidence suggested that transformation is necessary for downstream statistical analyses. Auto scaling emerged as the most suitable strategy for data scaling. Low thresholds for blank filtering were found to be the most effective in enhancing data quality. The optimization of filtering thresholds required a careful balance to remove unnecessary information while retaining vital data. This work provides an overview of commonly applied strategies in untargeted metabolomics analysis, emphasizing the importance of careful workflow selection and optimization. It serves as a resource for refining data processing strategies to achieve accurate and reliable results, while also offering fresh insights into the challenges encountered throughout the untargeted metabolomics processing pipeline. Created in BioRender. Aigensberger, M. (2024) https://BioRender.com/e63g016 . • A simple dataset generated from bovine saliva spiked with 42 compounds, analyzed by anion exchange chromatography and HR-MS. • Comparison of four optimized processing workflows based on XCMS, Compound Discoverer, MS-DIAL, and MZmine. • Manual classification of all detected features into true positive and false positive signals. • Comprehensive comparison of automatic and manual integration. • In-depth assessment of missing value imputation, transformation, scaling, and filtering.
Key Findings
1
Among overlapping features, MS-DIAL showed the greatest similarity to manual integration; XCMS and MZmine also performed well, while Compound Discoverer struggled with reliably integrating high-baseline peaks.
2
Autoscaling was identified as the most suitable data scaling strategy.
3
Comparison of four metabolomics workflows (XCMS, Compound Discoverer, MS-DIAL, MZmine) on a bovine saliva dataset spiked with 42 compounds revealed large differences in detected features.
4
Low thresholds for blank filtering improved data quality, but optimizing filtering thresholds requires balancing removal of noise and retention of vital data.
5
Missing values were mainly due to low-abundance signals, so imputation with small values was the most effective approach.
6
No clear evidence that data transformation is necessary for downstream statistical analyses.
7
Only approximately 8% of detected features were common to all four tools, indicating low overlap between peak tables.
Research Object
Untargeted metabolomics data processing workflows (XCMS, Compound Discoverer, MS-DIAL, MZmine) applied to bovine saliva LC-HRMS dataset spiked with small polar molecules
Research Subject
Comparison and evaluation of processing steps and post-processing strategies—feature detection and integration accuracy (overlap with manual integration), missing-value imputation, transformation, scaling, and filtering—and their impact on feature recovery and data quality
Publication Details
Publication Date
2024-11-27
Journal
Publisher
ISSN
Cited by
23
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest