Researchers at Friedrich Schiller University Jena developed an AI method to predict retention times of small molecules more reliably. Prof. Dr. Sebastian Böcker said, "The difficulty lies in the fact that retention times depend heavily on the experimental conditions." This new tool helps scientists identify complex substances in biological samples.
Whether in drug discovery, environmental analysis or metabolomics: anyone analyzing complex biological samples often needs to identify the small molecules they contain. Researchers at Friedrich Schiller University Jena, in collaboration with partners from the Helmholtz Zentrum München and the Technical University of Munich, have developed a method that addresses a problem in analytical chemistry that has persisted for decades.
The team, led by bioinformatician Prof. Dr. Sebastian Böcker, presents the new tool in Nature Methods.
An everyday problem in chemical analysis
Biological samples such as blood, cell material or bacterial cultures can contain a vast number of different molecules. While DNA and proteins can now be analyzed relatively well using established methods, the diversity of small molecules—known as metabolites—is particularly high. These include metabolic products, natural compounds, toxins, degradation products and many pharmaceutical substances.
Liquid chromatography is frequently used to separate such molecules from one another. In this process, a mixture of substances passes through a separation column; individual molecules are retained to varying degrees and therefore elute from the column at different times.
This so-called "retention time" provides an important indication of which substance is present in a sample. But when exactly does a particular molecule elute from the column? It is precisely this question that has occupied analytical chemistry for decades.
Why previous predictions have reached their limits
"The difficulty lies in the fact that retention times depend heavily on the experimental conditions: on the column used, the solvent, the gradient, the pH, the temperature and even on seemingly minor technical changes," says Böcker. "If, for example, a tube in the apparatus is replaced, or if a new tube is slightly longer than the old one, the measured times can shift significantly."
Previous models therefore often had to be trained or fine-tuned using data from the very same measurement system on which they were later to make predictions.
In practice, this means that researchers would first have to measure numerous standard substances before they could use the model effectively. This is time-consuming, expensive and, for many applications, simply not feasible.
First the order, then the time
The new approach from Böcker's team is not limited to a single system but can make predictions even for new systems and unknown molecules. The method focuses on the most commonly used type of liquid chromatography, the so-called "reversed-phase" mode, and operates in two steps. In the first step, a machine learning model calculates a so-called retention order index for a molecule.
This index does not directly describe a time but rather a molecule's position in the expected retention sequence. In the second step, this index is converted into specific retention times using a few known reference points.
Fleming Kretschmer played a key role in the development; he worked on the method as part of his doctorate and is a co-first author of the paper. "A key finding of our work is that our method outperforms other approaches that, unlike ours, first have to be trained extensively on the target system," emphasizes Kretschmer. "Our approach therefore enables precise predictions out of the box, even for new systems."
Relevance for drug discovery, natural product research and environmental analysis
The approach can help wherever unknown small molecules need to be identified. In natural product research, for example, researchers are looking for substances from bacteria, fungi or plants that could serve as new antibiotics, cancer drugs or active compounds against other illnesses.
In environmental analysis, food chemistry and pharmaceutical research, too, the same question arises time and again: What is contained in a complex sample—and does the measured retention time match the presumed chemical structure?
