Published in Nature Methods, the software tool addresses a decades-old analytical chemistry challenge by calculating a retention order index before mapping it to actual retention times using a small number of reference points.
The Challenge of Predicting Retention Times in Complex Biological Samples
Analyzing biological samples like blood, cell material, and bacterial cultures requires separating a vast array of small molecules known as metabolites. Liquid chromatography achieves this by passing a substance mixture through a separation column, where individual molecules exit at varying intervals called retention times. Bioinformatician Sebastian Böcker explains that these times fluctuate significantly based on experimental conditions, including the chosen column, solvent, pH level, temperature, and even minor hardware adjustments like replacing a length of tubing.
Because minor equipment shifts alter measurements, past predictive models required extensive training and tuning on the exact measurement system they would later use. This forced laboratories to measure numerous standard substances beforehand, making the process expensive and impractical for many routine analytical tasks.
How the Two-Step Machine Learning Tool Works
The newly developed approach eliminates the need to train models on specific target hardware by focusing on the most common reversed-phase liquid chromatography mode. Fleming Kretschmer explains that the method operates in two distinct phases:
- The algorithm calculates a retention order index, determining a molecule’s expected sequence position rather than an absolute time.
- The system applies a small set of known reference points to convert that ordinal index into concrete retention times.
This structure allows the software to deliver accurate, out-of-the-box predictions for unfamiliar systems and previously unmeasured molecules without prior system-specific calibration.
Applications Across Drug Discovery and Environmental Analysis
The software carries broad utility for identifying unknown small molecules across multiple scientific disciplines. In natural product research, investigators screen bacterial, fungal, and plant extracts for candidates that could serve as novel antibiotics or cancer treatments. Similar identification challenges arise in environmental analysis, food chemistry, and pharmaceutical development, where researchers must confirm whether observed retention times match suspected chemical structures.
The software package, designated as 2-step, is currently available as both a software package and a web application via the Böcker Lab GitHub repository. The development team anticipates that analytical laboratories will integrate the tool directly into existing mass spectrometry and liquid chromatography software suites.