Artificial Intelligence: Unexpected Results

Artificial intelligence (AI) is on the rise. Until now, AI applications generally have "black box" character: How AI arrives at its results remains hidden. Prof. Dr. Jürgen Bajorath, a cheminformatics scientist at the University of Bonn, and his team have developed a method that reveals how certain AI applications work in pharmaceutical research. The results are unexpected: the AI programs largely remembered known data and hardly learned specific chemical interactions when predicting drug potency. The results have now been published in Nature Machine Intelligence.

Which drug molecule is most effective? Researchers are feverishly searching for efficient active substances to combat diseases. These compounds often dock onto protein, which usually are enzymes or receptors that trigger a specific chain of physiological actions. In some cases, certain molecules are also intended to block undesirable reactions in the body - such as an excessive inflammatory response. Given the abundance of available chemical compounds, at a first glance this research is like searching for a needle in a haystack. Drug discovery therefore attempts to use scientific models to predict which molecules will best dock to the respective target protein and bind strongly. These potential drug candidates are then investigated in more detail in experimental studies.

Since the advance of AI, drug discovery research has also been increasingly using machine learning applications. As one "Graph neural networks" (GNNs) provide one of several opportunities for such applications. They are adapted to predict, for example, how strongly a certain molecule binds to a target protein. To this end, GNN models are trained with graphs that represent complexes formed between proteins and chemical compounds (ligands). Graphs generally consist of nodes representing objects and edges representing relationship between nodes. In graph representations of protein-ligand complexes, edges connect only protein or ligand nodes, representing their structures, respectively, or protein and ligand nodes, representing specific protein-ligand interactions.

"How GNNs arrive at their predictions is like a black box we can't glimpse into," says Prof. Dr. Jürgen Bajorath. The chemoinformatics researcher from the LIMES Institute at the University of Bonn, the Bonn-Aachen International Center for Information Technology (B-IT) and the Lamarr Institute for Machine Learning and Artificial Intelligence in Bonn, together with colleagues from Sapienza University in Rome, has analyzed in detail whether graph neural networks actually learn protein-ligand interactions to predict how strongly an active substance binds to a target protein.

How do the AI applications work?

The researchers analyzed a total of six different GNN architectures using their specially developed "EdgeSHAPer" method and a conceptually different methodology for comparison. These computer programs "screen" whether the GNNs learn the most important interactions between a compound and a protein and thereby predict the potency of the ligand, as intended and anticipated by researchers - or whether AI arrives at the predictions in other ways. "The GNNs are very dependent on the data they are trained with," says the first author of the study, PhD candidate Andrea Mastropietro from Sapienza University in Rome, who conducted a part of his doctoral research in Prof. Bajorath's group in Bonn.

The scientists trained the six GNNs with graphs extracted from structures of protein-ligand complexes, for which the mode of action and binding strength of the compounds to their target proteins was already known from experiments. The trained GNNs were then tested on other complexes. The subsequent EdgeSHAPer analysis then made it possible to understand how the GNNs generated apparently promising predictions.

"If the GNNs do what they are expected to, they need to learn the interactions between the compound and target protein and the predictions should be determined by prioritizing specific interactions," explains Prof. Bajorath. According to the research team's analyses, however, the six GNNs essentially failed to do so. Most GNNs only learned a few protein-drug interactions and mainly focused on the ligands. Bajorath: "To predict the binding strength of a molecule to a target protein, the models mainly 'remembered' chemically similar molecules that they encountered during training and their binding data, regardless of the target protein. These learned chemical similarities then essentially determined the predictions."

According to the scientists, this is largely reminiscent of the "Clever Hans effect". This effect refers to a horse that could apparently count. How often Hans tapped his hoof was supposed to indicate the result of a calculation. As it turned out later, however, the horse was not able to calculate at all, but deduced expected results from nuances in the facial expressions and gestures of his companion.

What do these findings mean for drug discovery research? "It is generally not tenable that GNNs learn chemical interactions between active substances and proteins," says the cheminformatics scientist. Their predictions are largely overrated because forecasts of equivalent quality can be made using chemical knowledge and simpler methods. However, the research also offers opportunities of AI. Two of the GNN examined models displayed a clear tendency to learn more interactions when the potency of test compounds increased. "It's worth taking a closer look here," says Bajorath. Perhaps these GNNs could be further improved in the desired direction through modified representations and training techniques. However, the assumption that physical quantities can be learned on the basis of molecular graphs should generally be treated with caution. "AI is not black magic," says Bajorath.

Even more light into the darkness of AI

In fact, he sees the previous open access publication of EdgeSHAPer and other specially developed analysis tools as promising approaches to shed light on the black box of AI models. His team's approach currently focuses on GNNs and new "chemical language models". "The development of methods for explaining predictions of complex models is an important area of AI research. There are also approaches for other network architectures such as language models that help to better understand how machine learning arrives at its results," says Bajorath. He expects that exciting things will soon also happen in the field of "Explainable AI" at the Lamarr Institute, where he is a PI and Chair of AI in the Life Sciences.

Mastropietro A, Pasculli G, Bajorath J.
Learning characteristics of graph neural networks predicting protein-ligand affinities.
Nat Mach Intell, 2023. doi: 10.1038/s42256-023-00756-9

Most Popular Now

MEDICA 2024 + COMPAMED 2024: Adapted Hal…

11 - 14 November 2024, Düsseldorf, Germany. The final preparations for MEDICA 2024 and COMPAMED 2024 in Düsseldorf have begun. A total of more than 5,500 exhibitors from approximately 70 countries...

AI does Not Necessarily Lead to more Eff…

The use of artificial intelligence (AI) in hospitals and patient care is steadily increasing. Especially in specialist areas with a high proportion of imaging, such as radiology, AI has long...

Commission Joins Forces with Venture Cap…

The Commission has launched a Trusted Investors Network bringing together a group of investors ready to co-invest in innovative deep-tech companies in Europe together with the EU. The Union's investment...

Why the NHS is Seeking to Make Media Ser…

Opinion Article by Dean Moody, Healthcare Services Director, Airwave Healthcare. Tim Kelsey and Martha Lane Fox called for WiFi to be made available free of charge throughout the NHS back in...

An AI-Powered Pipeline for Personalized …

Ludwig Cancer Research scientists have developed a full, start-to-finish computational pipeline that integrates multiple molecular and genetic analyses of tumors and the specific molecular targets of T cells and harnesses...

Philips and Medtronic Advocacy Partnersh…

Royal Philips (NYSE: PHG, AEX: PHIA), a global leader in health technology, and Medtronic Neurovascular, a leading innovator in neurovascular therapies, today announced a strategic advocacy partnership. Delivering timely stroke...

Wearable Cameras Allow AI to Detect Medi…

A team of researchers says it has developed the first wearable camera system that, with the help of artificial intelligence (AI), detects potential errors in medication delivery. In a test whose...

AI could Transform How Hospitals Produce…

A pilot study led by researchers at University of California San Diego School of Medicine found that advanced artificial intelligence (AI) could potentially lead to easier, faster and more efficient...

New AI Tool Predicts Protein-Protein Int…

Scientists from Cleveland Clinic and Cornell University have designed a publicly-available software and web database to break down barriers to identifying key protein-protein interactions to treat with medication. The computational tool...

Great Start for Ideas and Innovations: D…

8 - 10 April 2025, Berlin, Germany. From 15 October to 15 November 2024, the DMEA invites experts from business, science, politics and practice to actively participate in shaping the congress...

Start-Ups will Once Again Have a Starrin…

11 - 14 November 2024, Düsseldorf, Germany. The finalists in the 16th Healthcare Innovation World Cup and the 13th MEDICA START-UP COMPETITION have advanced from around 550 candidates based in 62...

AI for Real-Rime, Patient-Focused Insigh…

A picture may be worth a thousand words, but still... they both have a lot of work to do to catch up to BiomedGPT. Covered recently in the prestigious journal Nature...