Monday, 28 September 2026 PDT | 04:50 AM
The 1 News Alt Logo Text Smart News for Global Indians

Artificial Intelligence tool predicts stable hydrogen positions in drug-like molecules

AI News September 28, 2026 04:30 PM
Artificial Intelligence tool predicts stable hydrogen positions in drug-like molecules

A graph neural network trained on more than a million experimentally observed molecular structures can identify which of several possible chemical forms a drug-like molecule is most likely to adopt, addressing a longstanding gap in structure-based drug discovery

Researchers at New York University (NYU), New York, USA, have trained an artificial intelligence (AI) model to learn the chemical patterns associated with molecular stability and to predict – for drug-like molecules – where their hydrogen atoms are most likely to sit.

The work addresses the persistent problem in molecular design and drug discovery of being able to rapidly and reliably identifying the stable form of molecules that share a single molecular formula but can readily convert between related structures. Many drug-like molecules exist as two or more closely related forms – tautomers – in which a hydrogen atom moves from one site to another and the surrounding bonding pattern changes with it.

“Although this may seem like a small change, different tautomers of the same molecule can alter how a molecule interacts with a protein target,” said Dr. Yingkai Zhang, professor of chemistry at NYU and the study’s senior author.

“Correct tautomer assignment is consequently important for molecular modelling and structure-based drug discovery,” he added.

Identifying the correct tautomer has remained difficult, largely because of a scarcity of experimental data characterising tautomer structures. In the Protein Data Bank, the hydrogen positions that distinguish one tautomer from another are typically unavailable. The data it holds on structures are largely determined by X-ray crystallography and the resolution of macromolecular X-ray structures is generally insufficient to locate hydrogen atoms reliably, so their positions have to be inferred and consequently too a molecule’s tautomeric state.

Quantum mechanical methods are frequently too computationally demanding and expensive to apply across full compound libraries, while machine learning approaches have been constrained by the small size of experimentally characterised tautomer datasets in solution which tend to hold only a few hundred molecules.

By contrast, many high-resolution small-molecule X-ray crystal structures held in the Cambridge Structural Database, the world’s largest repository of experimental crystal structures, do show the position of hydrogen atoms.

“We realised that experimentally resolved hydrogen positions in high-resolution small-molecule crystal structures provide a largely untapped source of experimental information about tautomer stability,” said Zhang, who is also part of the NYU Simons Center for Computational Physical Chemistry.

Dr. Xiaolin Pan, a postdoctoral researcher in Zhang’s laboratory and the study’s first author, systematically mined the Cambridge Structural Database to build a dataset of more than 1.1 million tautomeric states which was orders of magnitude larger than existing experimental datasets.

The team then trained a graph neural network – a form of AI that uses deep learning to identify patterns between connected data points – to predict stable tautomers directly from two-dimensional molecular structures, without needing to have observed three-dimensional structures or carry out quantum mechanical calculations.

Applying the model to 5,075 PDBbind ligands – biomolecular complexes drawn from the Protein Data Bank – with multiple possible tautomeric states, the researchers identified 126 cases in which the tautomer originally assigned to the ligand was likely incorrect. This represented around 2.5 per cent of the database’s total. In each of these cases, the model proposed an alternative stable tautomer that showed improved hydrogen bonding patterns.

“Reassigning these tautomers generally produced more chemically reasonable interactions. This does not mean that the experimentally determined protein structures themselves are incorrect. Rather [that] our results suggest that the previously assigned chemical representation may warrant revision,” Zhang noted.

In one case, the model predicted a tautomer with a hydrogen atom positioned differently to the original assignment in the Protein Data Bank. As a result, the ligand formed additional hydrogen bonds with nearby protein residues, illustrating how a seemingly minor change can alter a molecule’s interactions with its surrounding protein environment.

More accurately identifying a molecule’s tautomer could help determine how it fits into a protein binding site during drug discovery and could improve the reliability of subsequent computer simulations. Molecular dynamics simulations, used to assess drug candidates, follow a molecule and its protein target over time to study how they interact and move.

“These calculations require the molecule’s hydrogen positions and chemical bonding to be correctly assigned and using the wrong tautomer could alter the predicted interactions and dynamics,” said Zhang.

The researchers have released their method as an open-source tool, Tautomer-Predictor, designed to rapidly analyse very large molecular libraries. In one test, it processed around 4.6 million compounds in 3.2 hours on a single GPU-enabled computing node.

For further reading please visit: 10.1039/d6sc03714c