A Perspective and Framework for Developing Sample Type Specific Databases for LC/MS-Based Clinical Metabolomics

Metabolites. 2019 Dec 21;10(1):8. doi: 10.3390/metabo10010008.

Abstract

Metabolomics has the potential to greatly impact biomedical research in areas such as biomarker discovery and understanding molecular mechanisms of disease. However, compound identification (ID) remains a major challenge in liquid chromatography mass spectrometry-based metabolomics. This is partly due to a lack of specificity in metabolomics databases. Though impressive in depth and breadth, the sheer magnitude of currently available databases is in part what makes them ineffective for many metabolomics studies. While still in pilot phases, our experience suggests that custom-built databases, developed using empirical data from specific sample types, can significantly improve confidence in IDs. While the concept of sample type specific databases (STSDBs) and spectral libraries is not entirely new, inclusion of unique descriptors such as detection frequency and quality scores, can be used to increase confidence in results. These features can be used alone to judge the quality of a database entry, or together to provide filtering capabilities. STSDBs rely on and build upon several available tools for compound ID and are therefore compatible with current compound ID strategies. Overall, STSDBs can potentially result in a new paradigm for translational metabolomics, whereby investigators confidently know the identity of compounds following a simple, single STSDB search.

Keywords: compound identification; database; metabolite identification; metabolomics; spectral library.