Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

RA-PLA: Retrieval-augmented graph convolutional networks for protein-ligand binding affinity prediction

  • Karim Abbasi ,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    karim.abbasi@khu.ac.ir, karim.abbasi@ut.ac.ir

    Affiliation Mosaheb Institute for Mathematical Research, Kharazmi University, Tehran, Iran

    ⨯
  • Hossein Banadkuki,

    Roles Conceptualization, Funding acquisition, Resources, Validation, Writing – original draft, Writing – review & editing

    Affiliation Mosaheb Institute for Mathematical Research, Kharazmi University, Tehran, Iran

    ⨯
  • Reza Sepahvand,

    Roles Conceptualization, Data curation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Computer Engineering and Information Technology Department, University of Qom, Qom, Iran

    ⨯
  • Shima Rashidi

    Roles Conceptualization, Formal analysis, Validation, Visualization, Writing – review & editing

    Affiliation Department of Artificial Intelligence and Data Science, College of Science and Technology, University of Human Development, Sulaymaniyah, Kurdistan Region of Iraq

    ⨯

Abstract

Protein-ligand binding affinity (PLA) prediction is essential for computational drug discovery, yet existing deep learning methods typically employ a single globally trained model that fails to adapt to individual query samples. In this paper, we propose a novel end-to-end framework that retrieves hard protein-ligand pairs—samples whose nearest neighbors have conflicting labels under manifold smoothness constraints—and integrates them via a semi-supervised graph convolutional network (GCN). For each query, we automatically construct a graph where nodes represent protein-ligand pairs and edges encode pairwise similarities. Our model jointly learns joint descriptors, graph topology, and the GCN predictor. During inference, we fine-tune the model per query using its retrieved hard neighbors.We evaluate on four benchmarks: PDBbind, Davis, KIBA, and BindingDB. Our method consistently outperforms state-of-the-art approaches. On Davis, we achieve a Concordance Index (CI) of 0.952 ± 0.005 (1.6% improvement over NerLTR-DTA) and AUPR of 0.815 ± 0.005 (7.6% improvement). On KIBA, we achieve CI of 0.916 ± 0.002 and AUPR of 0.871 ± 0.005, outperforming DeepCDA by 2.7% and 5.9%, respectively. On PDBbind, our CI of 0.894 ± 0.002 represents an 11.4% improvement over PLA-MoRe. On BindingDB, we achieve AUPR of 0.518 ± 0.009 (5.9% improvement over DeepCDA). In cold-target generalization across four unseen protein families, we achieve average AUPR improvements of 15.5%. Paired t-tests confirm statistical significance (p < 0.05), and ablation studies validate that both hard sample retrieval and end-to-end learning are essential for performance gains. Our framework demonstrates that query-specific adaptation significantly enhances both accuracy and generalization in PLA prediction.

1. Introduction

Protein-ligand binding affinity (PLA) prediction is a crucial step in the computational drug discovery process [1–3]. The primary goal of PLA is to determine the binding affinity between a drug (ligand) candidate and a protein sequence [4,5]. PLA measurement through experiments is expensive and time-consuming. Hence, computational-based methods receive more attention [6–8]. To date, numerous approaches have been proposed in this research area. Computational approaches generally consist of two steps: feature extraction and task prediction. In the feature extraction step, discriminative features are extracted from the raw input. Feature extraction methods are divided into data-driven and non-data-driven methods [9–12]. In data-driven methods, features are learned automatically, while in non-data-driven approaches, features are extracted manually (hand-crafted features). The task prediction step takes the features as input and then maps them to the interaction space.

In classic PLA prediction approaches, features are usually extracted through non-data-driven methods like extended-connectivity fingerprints (ECFP) [13], rapid overlay of chemical structures (ROCS) [14], molecular linear notation by circular traverse (MLNCT), and fragment-based descriptors [15,16]. Many hand-crafted feature extraction algorithms are available [3,17]; hence, one of the main challenges in this area is determining which feature extraction algorithms are appropriate for the specific task.

In recent years, deep learning-based models have received more attention and improved the performance of PLA prediction [18–20]. In most deep learning-based approaches, first, a feature encoder for drug molecules and a feature encoder for protein sequences are trained. Then, these feature vectors are fused to construct a feature representation for the protein-ligand pair, which is then fed into the MLP to predict the affinity value. In other words, a single learned model can be used to predict the affinity value for all unseen protein-ligand pairs. We believe that the generalizability of this approach is limited. For each unseen protein-ligand pair, some training protein-ligand pairs might have more information than others. To incorporate this knowledge, we have designed a model that retrieves some training protein-ligand pairs for each unseen protein-ligand pair and uses those pairs to learn a specific model for each unseen pair. It means that each test sample has its prediction model, which helps the model improve its performance.

To incorporate the knowledge of the retrieved training samples in inferring the label of the query sample, we have utilized a semi-supervised learning method. The reason is that all of the retrieved samples are labeled training samples, and the query sample is unlabeled (semi-supervised setting). We have chosen to utilize semi-supervised GCN as the learning method to infer the label of the query sample. In this case, the nodes of the graph correspond to the feature vectors of the protein-ligand pairs, and the edges between the nodes represent the similarity between protein-ligand pairs. Based on this setting, to infer the label of the query sample, GCN leverages the knowledge of neighborhood protein-ligand pairs (i.e., similar pairs), which aligns with the paper's motivation.

To design a model with high generalization ability, we develop a robust task prediction network model by integrating hard similar samples into the test sample. Hard protein-ligand pairs are retrieved and integrated into interaction affinity prediction using a semi-supervised graph convolutional network for each test sample. To determine hard samples, the similarity of the query sample is computed with all training samples and then sorted. In the sorted vector, a sample is labeled as hard if its nearest neighbors have differing class labels. It should be noted that the class labels refer to the active/inactive classes.

To recap, the proposed model is a unified end-to-end framework with three specialized sub-networks: 1) Descriptor Sub-network: learns joint representations for protein-ligand pairs, 2) Adjacency Sub-network: learns pairwise similarities to construct the graph topology, and 3) GCN Prediction Sub-network: propagates information through the graph to predict test sample affinity. This work is the first to demonstrate that modeling the graph topology between test samples and their hard training neighbors improves prediction performance. All components are optimized jointly during end-to-end training.

The paper is structured as follows: Section 2 reviews related work on affinity prediction and graph-based learning. Section 3 details the datasets, while Section 4 introduces the proposed unified framework. Finally, Section 5 evaluates the model through benchmarks and ablation studies, and Section 6 discusses results, limitations, and future directions.

2. Related work

In this section, the recent advances in PLA prediction are reviewed. Ozturk et al.[21] utilize a CNN to extract features for both compound and protein. Abbasi et al. [17] introduced an approach that utilizes CNN, LSTM, and a two-sided attention mechanism to extract relevant features. Also, they employed domain adaptation techniques to learn more about discriminant features. These approaches use fully connected layers as prediction layers. Some methods utilize auxiliary knowledge, such as protein-protein interaction (PPI) graphs, to design a more accurate PLA prediction model [22–24]. Wang et al. [25] applied the random walk with restart method to calculate the topological similarity matrices using the PPI network, gene ontology, and drug fingerprint, as well as drug side effects. They utilized a multimodal deep autoencoder (MDA) to integrate multiple similarities and learn high-level features, which were then fed into a multi-layered neural network to predict ligand and protein interactions. Lim et al. [26] introduced a distance-aware graph attention algorithm to extract graph features of intermolecular interactions from 3D structural information. The learned features are fed into the fully connected layers to predict interaction affinity between the ligand and protein target. Ru et al. [27] introduced a method that utilized a ranking framework to predict affinity values, called NerLTR-DTA. They used learning-to-rank (LTR) algorithms to learn the model. Their approach provides the priority order of proteins related to the query drug and the priority order of ligands to the target protein. Zhou et al. [7] proposed a multimodal drug-target interaction prediction model called MultiDTI. Their method integrates the heterogeneous network into drug-target interaction by mapping it into a common space. This is done by integrating drug similarity, target similarity, side effects, and disease into the loss function. Shim et al. [28] used the similarity of drug-drug and protein-protein interactions to provide a 2D input to feed into the 2D convolutional layer. The corresponding columns in the drug-drug and protein-protein similarity matrix are taken for a drug-protein pair, and its outer product is computed. Yuan et al. [22] concentrated on feature aggregation and introduced a FusionDTA approach. They introduced a novel multi-head linear attention mechanism to aggregate global information. Aleb [29] has introduced two multilevel attention-based models for PLA prediction. Nguyen et al. [30] introduced a method in which a drug is represented as a graph. Then, it is fed into a graph convolution network (GCN), graph attention network (GAT), graph isomorphism Network (GIN), and GAT-GCN to learn the representation of the drug molecule. To represent the protein sequence, they used a 1D-CNN. They have shown that representing drugs as graphs could improve performance. Son and Kim [31] introduced a method that utilized a graph convolutional network called GraphBAR. In their approach, the ligand-protein complex is fed as input to GCN. Each graph has multiple adjacency matrices, which are created using different distance metrics. Liu et al. [32] introduced a new molecular representation in which ligand-protein complexes are modeled as Dowker complexes. Then, combinatorial Laplacian matrices are constructed by generating at different scales and passed into the Riemann zeta functions to generate molecular descriptors. Nguyen and Wei [33] introduced a novel algebraic graph learning score (AGL-Score) model to generate low-dimensional representations from physical and biological information. Li et al. [34] introduced a protein-ligand binding affinity prediction model called PLAMoRe. They proposed a new representation of learning that utilizes both structural knowledge and bioactive properties. Autoencoders are used to learn bioactive feature encoders from multisource information, including chemical, target, network, cellular, and clinical data. Wang et al. [35] introduced a new deep learning-based protein–ligand binding affinity prediction approach named DeepDTAF. In their approach, the protein-binding pocket is first used as input to extract the local feature, which is then combined with the global context feature. They have used a dilated convolutional network to extract long-range interactions.

3. Datasets

In this section, datasets are explained. To evaluate the proposed approach, four well-known datasets, including PDBbind [36], Davis [37,38], KIBA [39,40], and BindingDB [41], are used. The following sections provide detailed explanations of each dataset.

PDBbind. We have utilized PDBbind v2016, which contains three sets: the general set, refined set, and core set. The general set is filtered to select better-quality ligand-protein complexes, called the refined set. In this dataset, proteins are recorded in pdb format, and ligand structure data are recorded in mol2 or sdf format. We have chosen training, validation, and test sets that are consistent with other methods to ensure a fair comparison. Therefore, the core set is used as a test set [42].

Davis. This dataset comprises the interactions of 442 unique proteins (from the kinase protein family) and 68 unique ligands. The respective dissociation constant (Kd) value measures the binding affinity of protein-ligand interaction. The Kd measure is transformed into log space (𝑝K𝑑), similar to [17,37], as follows:

(1)

KIBA. This dataset includes 246,088 interactions between 467 unique proteins and 52,498 unique ligands. This dataset is filtered, so proteins or ligands with fewer than ten interactions are removed. This dataset measures binding affinity using the KIBA score, which combines kinase inhibitor bioactivities from various sources, including IC50, Ki, and Kd [43].

BindingDB. It is a public dataset that provides the binding affinity of each pair, stated using at least one measure, such as IC50, EC50, Ki, and Kd. In this study, the Ki-labelled samples, including 234491 pairs, are used. To evaluate the generalizability of the approach, similar to [17,44], we exclude the four classes of proteins entirely, including nuclear estrogen receptors (516 samples), ion channels (8,101 samples), receptor tyrosine kinases (3,355 samples), and G-protein-coupled receptors (77,994 samples) from the training set. The remaining protein-ligand pairs are divided into the training set with 101,134 samples and the test set with 43,391 samples.

4. Method

In this section, the proposed model is explained in detail. The overall schematic of the proposed approach is presented in Fig 1. The following section begins with a presentation of the problem formulation. Then, feature extraction and task prediction steps are explained.

thumbnail
Fig 1. The overall schematic of the proposed approach using Strategy 3 (hard sample retrieval with dropout).

First, hard samples (those with conflicting neighbor labels) are retrieved for the query, combined with 50% random dropout selection. Then, the query sample and all retrieved samples are fed into the model to predict the affinity value. It should be noted that * means that the feature extraction module is shared between the retrieval procedure and the forward pass of the model.

https://doi.org/10.1371/journal.pone.0357278.g001

4.1. Problem formulation

Let c and p, respectively, denote a drug compound (ligand) and a 1D protein sequence of amino acids. All training samples are shown by where N denotes the number of training samples; it should be noted that each sample is a protein-ligand pair. Also, shows the strength of binding, which can be either a binary binding affinity (active or inactive) or a quantitative binding affinity (e.g., IC50, Kd, Ki). In the proposed approach, both cases are considered.

In the test step, a test sample is shown by is fed into the proposed model. To this end, in the first step, a set of hard training samples for the test sample is retrieved, as shown by , where indicates the number of retrieved similar training samples. It should be noted that the number of retrieved samples might differ for each input query sample (see section 4.2). The main goal is to use this knowledge to predict the binding affinity of the test sample.

4.2. Retrieval step

This step retrieves samples from the training dataset for each input query sample. Ground truth knowledge could help to retrieve samples during the training phase. However, this knowledge is not provided in the test phase. We have suggested three ways to retrieve samples. In the first approach, we could compute the distance between the feature descriptors of the test samples and all the training samples using a distance metric, such as the Euclidean distance metric, and then retrieve the closest ones. The disadvantage of this step is that if the similarity measure and feature descriptor fail to account for the distance between samples, this error is propagated to subsequent steps. The second approach is to retrieve samples randomly such that half of the retrieved samples belong to the active samples class and the other half belong to the inactive class. Half the samples were retrieved randomly. The third approach is to retrieve samples automatically. To clarify, the third approach involves a multi-step process. First, we compute the Euclidean distance between the query sample's feature descriptor and all positive samples (i.e., active samples) to measure similarity. Next, we rank the positive samples based on these similarity values. For the top half of the retrieved samples, we analyze each sample's neighborhood. If a sample has neighboring samples with conflicting labels, we flag it as a “hard sample” and include it in our retrieval set. These hard samples are particularly valuable because they tend to lie near the decision boundary, where model uncertainty is highest, making them critical for refining the model's discriminative ability.

The retrieval step identifies hard training samples for each query. Let denote the training set, where is the feature descriptor of protein-ligand pair (learned by the feature encoder) and is its continuous binding affinity. For binary classification, we define as the binary label obtained by thresholding (thresholds described in Section 4).

For a query with feature descriptor , we compute the Euclidean distance to all training samples:

(2)

The k-nearest neighbor set is defined as:

(3)

In all experiments, we set .

For practical comparison with existing methods, we also employ a binary version:

(4)

That is, a sample is hard if its binary label conflicts with at least one of its nearest neighbors, indicating it lies near the decision boundary.

The hard sample retrieval set is then:

(5)

The size of is variable and depends on the local structure of the feature space around .

We define hard samples based on the principle of manifold smoothness in feature space. The key assumption is that samples with similar molecular structures (and thus similar feature representations) should have similar binding affinity values. Hard samples are those that violate this assumption—they are close in feature space but have dissimilar outputs.

For practical purposes, we use a binary approximation where continuous affinity values are thresholded to obtain active/inactive labels. A training sample is considered hard if it is among the nearest neighbors of the query in feature space but has a different binary label than the query. This identifies samples that lie near the decision boundary in feature space—regions where the structure-affinity mapping is challenging and where model uncertainty is highest.

For samples that don't meet the hard sample criteria, we apply random dropout selection with a 50% probability to maintain diversity in our training set. Finally, we repeat this entire process for inactive samples to ensure balanced selection. This approach enables us to focus on challenging and informative examples while preventing the model from overlooking less obvious but still important cases.

Regarding the disadvantage of the first approach, we acknowledge that the similarity measure and feature descriptor may fail to account for the distance between samples, leading to errors in the retrieval process. The third approach overcomes this limitation in two ways:

  • Robustness to noise: By using a neighborhood analysis, we can identify hard samples that are likely to be noisy or outliers, and add them to the retrieval set. This helps to reduce the impact of noise on the retrieval process.
  • Diversity in retrieval: The dropout selection mechanism ensures that we have a diverse set of samples in the retrieval set, which helps to avoid bias in the results. By randomly selecting samples with a predetermined probability, we can explore new options and avoid overfitting to a specific subset of samples.

We believe that the third approach strikes a good balance between exploration and exploitation, thereby enhancing the model's efficiency. By focusing on hard samples near the decision boundary, we can improve the model's discriminative power and robustness. Meanwhile, the dropout selection mechanism helps maintain diversity and prevent overfitting.

The intuition behind the proposed retrieval approach is based on manifold smoothness; the local neighborhood around a given query point should belong to the same class label. If we could train a model that considers this concept, we could improve the model's generalization ability. We could train a more discriminative and generalizable model by ensuring local manifold smoothness around labeled and unlabeled data. In the proposed approach, this issue is considered by retrieving samples that do not satisfy the manifold smoothness and utilizing these samples in the task prediction step (through semi-supervised GCN).

4.3. Feature encoder

This study uses a convolutional neural network (CNN) to design a feature encoder for protein-ligand pairs. For this purpose, two 1D CNNs are considered for ligand and protein sequences (see Fig 1, feature extraction module). Let and denote the obtained feature vector of drug and protein CNNs, respectively, where and respectively show the length of the SMILES of a drug molecule and the length of the protein sequence. The drug feature extraction network outputs a matrix where each row represents the feature vector of the corresponding SMILES token in the input sequence. Similarly, the protein feature extraction network generates a matrix in which the row encodes the feature vector of the amino acid in the input protein sequence. Then, the outputs of these CNNs are concatenated to produce the descriptor of the protein-ligand pair, which is shown by . In the proposed approach, the CNN architecture is designed similarly to [17].

4.4. Semi-supervised GCN

The prediction step is performed using a semi-supervised graph convolutional network. This network comprises three hidden graph convolutional layers (see Fig 2). The input and output of each graph convolutional layer are the graphs. Hence, we need to construct a graph to feed into the GCN. In the proposed approach, each graph node corresponds to a protein-ligand pair. The graph has nodes, where one of these is a test protein-ligand pair, and the remaining ones are the retrieved training protein-ligand pairs. Let X and A denote the node descriptor matrix and adjacency matrix of the input to GCN. Each row of the matrix X represents a node, a descriptor of a protein-ligand pair in our approach. The output of the feature encoder network is used to construct the node descriptor matrix X. The Adjacency matrix is a square matrix where each matrix element indicates whether the corresponding protein-ligand pairs are adjacent or not. It should be noted that out of nodes of the graph correspond to the training samples (retrieved samples for the query pair); therefore, we could utilize the label information of these samples to construct some part of the matrix . Hence, elements of matrix A are shown by . These elements are set to one if the corresponding nodes have similar labels. Otherwise, they are set to zero. The other elements (connections between test samples and training samples) are learned during the training phase, which is explained in section 4.5.

thumbnail
Fig 2. The proposed approach's overall schematic of the task predictor network.

After constructing a graph using the query sample and retrieved samples, it is fed into the GCN to predict the affinity value of the query sample.

https://doi.org/10.1371/journal.pone.0357278.g002

A GCN processes the constructed graph to extract structural features, which are then passed through a Multilayer Perceptron (MLP) to predict protein-ligand binding affinity. For the regression task, we use a linear activation function in the final layer and optimize using mean squared error loss:

(6)

where is the predicted label for the ith sample. For binary classification of binding activity (active/inactive), we employ categorical cross-entropy loss.

4.5. Graph convolutional network

This section explains how a graph convolutional network works. The following describes the function of the graph convolutional layer. It takes X and A as input and then computes the output with m filters as follows:

(7)

where is a matrix of filter parameters, and is an activation function. The descriptions of matrices X and A are given in the previous section. Also, , where is an × identity matrix. The elements of are defined as follows:

(8)

In the case of the classification task, the activation function of the last layer is the sigmoid function. From a biological perspective, in our approach, when a graph passes through a graph convolution layer, the feature vectors of each node's corresponding neighboring nodes (similar protein-ligand pairs) are combined to form a new feature vector for the node. In these combinations, only similar samples are incorporated in learning a new feature vector for the protein-ligand pair. Hence, the adjacency matrix plays a crucial role because it defines the neighboring nodes. The loss function of this network is defined as follows:

(9)

Where c is the number of the class label, and denotes the output of the GCN for the ith training protein-ligand pair.

4.6. Adjacency matrix

In this section, the adjacency of the input sample and training samples are learned. To do so, the network is designed. In the proposed approach, two samples are considered similar if they have the same label. Hence, we predict the label for the input sample and then use the predicted labels to compute the adjacency of the test sample and training sample. To do so, the network includes a self-attention layer and a multi-layer perceptron (MLP) (see Fig 1- learn adjacency matrix module). The attention layer computes the interaction of the input (feature vector) with itself to determine which part of the input should pay more attention to. The outputs are aggregations of these interactions and the attention scores. Then, the attention layer output is fed into the MLP, which contains multiple fully connected layers. The last layer is a fully connected layer with one neuron whose activation function is sigmoid.

The main goal of this layer is to predict the label of the input sample to utilize it to compute its adjacency with retrieved training samples. They are similar if the input and retrieved training samples have the same label. Network performs as a binary classifier: active or inactive. All training data is used to train this network. If the labels of samples are continuous, they should be converted to discrete through thresholding. Let denote the output of the network . In this case, the loss function of the network is a binary cross-entropy, which is defined as follows:

(10)

where is the ground-truth label of the ith protein-ligand pair. If is continuous, it is converted to a binary class (active/inactive) through thresholding. Also, denotes the output of the network.

Given an input query, which consists of a drug-protein pair shown by , we first pass it through the feature extractor network to obtain a set of extracted features. These features are then used as input to the module, which predicts a pseudo label for the input pair as follows:

(11)

In parallel, we compute the set of retrieved samples that are most relevant to the input pair. To construct the graph that fed into the GCN, we need to define the adjacency matrix, which represents the connection knowledge between the input pair and the retrieved pairs as follows:

(12)

Where is a vector with dimension, which denotes the labels of the retrieved samples. Also, is the element-wise product operator. The repmat operator replicates an array () several times (in this formula, times). By putting and together, the adjacency matrix A is created. The following preprocessing should be done to fed it as an adjacency matrix to the GCN layer:

(13)(14)(15)

These formulas are used to normalize a graph's adjacency matrix, which is a common step in graph analysis and machine learning. The process involves modifying the matrix, calculating node degrees, and applying a normalization technique to scale the matrix. The resulting normalized matrix can be used in various applications, such as graph clustering, node embedding, and network analysis. The matrix A is finally fed into the semi-GCN as the adjacency matrix.

So far, we've established two key loss functions. The primary loss function, , concentrates on optimizing the overall output of our proposed model. In contrast, the secondary loss function, , targets the intermediate output, which plays a crucial role in forming the adjacency matrix. The overall loss function of the proposed approach is . The proposed approach is learned end-to-end. To learn the proposed network, at first, the feature encoder network along with the adjacency network is trained using training data . Then, we learn a semi-supervised GCN using the constructed graph of the input sample and the retrieved training samples.

5. Method evaluation

In this section, experimental results are given. The proposed approach is evaluated on four well-known datasets in PLA, including PDBbind, Davis, KIBA, and BindingDB. A detailed description of these datasets is provided in Section 3.

To compare the proposed approach with the current state-of-the-art approaches, we choose KronRLS [45], SimBoost [36], DeepAffinity [44], DeepDTA [21], DeepCDA [17], SimCNN-DTA [28], GraphDTA [30], MSF-DTA [46], GEFormerDTA [47], TEFDTA [48], and NerLTR-DTA [27] as baseline approaches. In this case, the results of base approaches are taken directly from their original publications, except that we explicitly state that we train them.

In the proposed approach, the nested cross-validation approach is similar to [17,21,37], is used on the KIBA and Davis datasets. In this case, the dataset is randomly split into six equal parts, where five parts are chosen as the training set via 5-fold cross-validation, and the remaining part is used as the test set. It should be noted that the test set is entirely different from the training sets. The final results are the average of the five outcomes. In BindingDB, the training and test sets are obtained similarly to [17,44]. It should be noted that in whole approaches, we use the same split with comparable methods for a fair comparison.

This study evaluates the proposed approach in two ways: continuous binding affinity prediction (a regression task) and binary binding affinity prediction (a classification task). Since continuous binding affinity is provided in the four mentioned datasets, we use thresholds to convert it to binary (active or inactive) values. To this end, we have done something similar to [17,37]. The thresholds for the -logkd, KIBA, and -logki scores are 7, 12.1, and 7.6. In this case, all datasets, including both active and inactive samples, exhibit an uneven distribution. The unit of all measures, including ki, KIBA, and kd is nm/L, where nm denotes nanomole. Note that ki and kd are transformed to the log space. For the PDBbind dataset, there is no predefined threshold to convert the continuous label to a discrete one. Hence, we get the median of the labels of the training samples and use it as the threshold.

For regression tasks, Concordance Index (CI), Mean Squared Error (MSE), index and the Pearson correlation coefficient (R) are used as evaluation metrics. CI investigates the concordance between the ground truth value and the predicted value. The intuition behind CI is that two protein-ligand pairs with different binding affinity values should be in the correct order. MSE is a well-known measure, and the index considers how the predictive value tends toward the mean. The Pearson correlation coefficient (R) indicates the linear correlation between the predicted and the ground truth values. If the predicted value and the ground truth value have a strong correlation, the Pearson correlation coefficient (R) gets a value of 1. Also, for the classification task, the Area Under the Precision-Recall curve (AUPR) is used as an evaluation measure [49]. As mentioned, in binary binding affinity, all datasets are imbalanced. Hence, AUPR is selected because it is suitable for situations with a significant skew in the class distribution.

The hyper-parameters optimization is done for the number of the filters (same for proteins and ligands) search over [16,32,64, 128, 256], the length of the filter size for ligands over [4,6,8,12], the length of the filter size for proteins over [8,12,16], and the learning rate for training network search over [0.1,0.01,0.001, 0.001,0.0005,0.0001,0.00005,0.00001]. The filter size for the protein sequence is larger than the SMILES sequence because the protein sequence length is greater than the length of the SMILES sequence. Additionally, to strike a balance between model performance and efficiency, we refrain from using large filter sizes. The results show that in most cases, eight for the ligand filter size and 12 for the protein filter size led to better performance. Also, 0.00001 is a more reliable learning rate for training the model. For the retrieval step, all the training samples are desired to be utilized in graph construction; however, this is computationally intractable.

This paper runs the program on a computer with an Intel(R) Core (T.M.) i9 CPU, NVIDIA GeForce GTX 1070 8 GB GDDR5, and 32 GB DDR4 RAM. We implemented our method using Python 3.8, TensorFlow, and Keras.

5.1. Ablation study

To evaluate the effectiveness of each module and key hyperparameters in the proposed method, we conducted ablation studies with two complementary analyses: First, we created three architectural versions. The first version, shown by , we only utilized the drug and protein feature extractors to extract features from the input data. The concatenated features were then fed into a multi-layer perceptron (MLP) to predict the binding affinity. This version did not utilize semi-supervised GCN and is similar to the DeepDTA approach. The second version, shown by , we used a first model to produce features for the drug and protein, which were then fed into a semi-supervised graph convolutional network (GCN). However, the learning of the drug and protein feature encoders was completely independent of the semi-supervised GCN, making this approach not end-to-end. The third version, called the full proposed model, includes the drug and protein feature extractors, as well as the semi-supervised GCN.

The results of the first ablation study on Davis are given in Table 1. As shown, the whole proposed model (the third version) achieves the best performance, outperforming the other two versions. This suggests that the knowledge retrieved from the samples and the end-to-end learning of the drug and protein feature encoders are essential components of the proposed method.

thumbnail
Table 1. Ablation study results on the Davis dataset evaluating the impact of different components of the proposed approach.

https://doi.org/10.1371/journal.pone.0357278.t001

The results of the ablation study demonstrate the effectiveness of each module in the proposed method. The use of knowledge from retrieved samples and the end-to-end learning of the drug and protein feature encoders significantly improves the model's performance. The two-stage approach (the second version) performs worse than the complete proposed model, indicating that the independent learning of the drug and protein feature encoders is less effective than the end-to-end learning approach. The version with only drug and protein feature extractors () performs the worst, highlighting the importance of utilizing the knowledge from the retrieved samples and the semi-supervised GCN in the proposed method.

Second, we systematically investigated the impact of retrieval set size on model performance, as this hyperparameter critically balances two key objectives: exploring diverse samples through random selection and exploiting informative cases through hard sample selection. We evaluated various retrieval set sizes {50, 100, 150, 200, 300} and found optimal performance at 150 samples. The results obtained are presented in Table 2. As shown, when the number of retrieved samples is set to 150, the proposed approach achieves better performance. Through extensive empirical validation, we observed that smaller retrieval sets may limit the model's exposure to challenging cases, while giant sets can dilute the importance of informative hard samples. Our chosen threshold of 0.50 for random selection – applied to both positive and negative samples- was carefully determined through systematic experimentation to optimize this trade-off.

thumbnail
Table 2. Performance comparison using different retrieval set sizes (50, 100, 150, 200, 300) on the Davis dataset.

https://doi.org/10.1371/journal.pone.0357278.t002

Third, to justify our selection of the third retrieval strategy (hard sample retrieval with dropout), we evaluated all three strategies described in Section 4.2 on the Davis dataset. The results are presented in Table 3.

thumbnail
Table 3. Comparison of retrieval strategies on Davis dataset.

https://doi.org/10.1371/journal.pone.0357278.t003

  • Strategy 1 (Distance-based): Retrieved the k-nearest training samples based on Euclidean distance in the feature space.
  • Strategy 2 (Random balanced): Retrieved equal numbers of active and inactive samples randomly.
  • Strategy 3 (Hard sample + dropout): Retrieved hard samples (those with conflicting neighbor labels) combined with 50% random dropout selection.

As shown in Table 3, the third strategy consistently outperforms both alternatives across all metrics. Strategy 1 suffers from propagation of errors in similarity measurement and feature representation, while Strategy 2 lacks the informative hard samples that lie near decision boundaries. Strategy 3 strikes the optimal balance between exploration (diverse samples) and exploitation (challenging, informative cases).

To evaluate the contribution of our label-dependent adjacency construction, we compared it with a standard k-NN (label-free) adjacency strategy. In the k-NN strategy, edges are established based solely on feature distance with equal weights, while our approach uses feature distance pseudo-labels for edge existence.

As shown in Table 4, our approach significantly outperforms k-NN (k is set to 3) across all metrics (CI: 0.952 vs. 0.918, p < 0.01). The improvement is particularly notable for AUPR (0.815 vs. 0.786), suggesting that label-dependent edge refinement is especially beneficial for distinguishing active from inactive compounds.

thumbnail
Table 4. Comparison of k-NN (label-free) adjacency and our proposed label-dependent adjacency strategy on the Davis dataset.

https://doi.org/10.1371/journal.pone.0357278.t004

These results confirm that incorporating label information into adjacency construction—when properly constrained by feature-based connectivity—provides meaningful improvements over purely feature-based adjacency.

The comparison with k-NN adjacency (Table 4) demonstrates that our approach substantially outperforms label-free feature-based adjacency (CI: 0.952 vs. 0.918), confirming that the improvement comes from meaningful graph refinement rather than overfitting to label patterns.

5.2. Results

This section compares results on the PDBbind, Davis, KIBA, and BindingDB datasets with the base approaches.

For a fair comparison, the comparison should be conducted under the same setting as the state-of-the-art approaches. In PDBbind, two different settings are employed in state-of-the-art techniques. The first setting is the refined set as the core set (285 pairs), and the remaining samples are used as training samples. The results obtained in this setting are presented in Table 5. In this case, the proposed approach improves the Pearson correlation coefficient by 7.3% over the 1D2D-CNN approach.

thumbnail
Table 5. Comparing the proposed approach results with base approaches on the PDBbind Dataset (the first setting).

https://doi.org/10.1371/journal.pone.0357278.t005

In the second setting, PDBbind is divided randomly into six parts, where one part is designated as the test set, and the remaining five parts are used to train the model using five-fold cross-validation. The results obtained using the PDBbind dataset in the second setting are presented in Table 6. As shown, the proposed approach outperforms the other approaches significantly. It improves the CI measure by 11.4% compared to the PLA-MoRe approach. Also, in MSE, our approach gets the lowest value compared to the other methods. One of the main contributions that led to better performance is that the knowledge of the retrieved samples is also considered in the task prediction. Based on the retrieval algorithm, all hard samples, neighboring samples near the query sample but with different labels, are incorporated to infer the label of the query sample. The class labels refer to the active/inactive classes.

thumbnail
Table 6. Comparing the proposed approach results with base approaches on the PDBbind Dataset (the second setting).

https://doi.org/10.1371/journal.pone.0357278.t006

Regarding the datasets in which the corresponding label for each protein-ligand pair is a continuous value, we should convert the continuous value to a binary value using thresholding. However, there is no predefined threshold for the PDBbind dataset to convert the continuous label to a discrete one. Hence, we get the median of the labels of the training samples and use it as the threshold. It should be noted that the retrieved hard samples serve as support vectors in SVM, aiding the model in learning more accurate decision boundaries for the query sample.

Also, in this paper, for evaluation, we use three external test sets: the PDBbind 2013 core set, the PDBbind 2016 core set, and the PDBbind 2019 holdout set. The 2019 holdout set simulates a temporal split scenario, where a model trained on prior structures is used to predict binding affinities for later released structures. Notably, there is no overlap among training, validation, and testing sets.

To ensure fair comparison and to provide a comprehensive evaluation across all standard test versions, we report results on all three test set versions in Table 7. Our training set is consistent with the protocol in [32], which has been widely adopted for evaluating structure-based affinity prediction methods.

thumbnail
Table 7. Comparative evaluation of RA-PLA on three standard PDBbind test sets (2013 core set, 2016 core set, and 2019 holdout set).

https://doi.org/10.1371/journal.pone.0357278.t007

The comparative evaluation of RA-PLA across the three PDBbind test sets reveals a clear and expected performance gradient, with the model achieving its strongest results on the 2013 and 2016 core sets and a notable decline on the more challenging 2019 holdout set. Specifically, the concordance index (CI) remains robust for the 2013 and 2016 sets at 0.885 and 0.894, respectively, alongside low mean squared errors (MSE) of 1.019 and 0.961 and high squared correlation coefficients () of 0.681 and 0.674, indicating reliable predictive power and good generalizability for these earlier benchmarks. However, performance drops significantly on the 2019 set, with the CI falling to 0.812, the MSE nearly doubling to 1.456, and the decreasing to 0.619, which suggests that RA-PLA struggles with the greater chemical and structural diversity or the more difficult binding affinity cases present in the newer data. This degradation highlights a key limitation in the model's ability to extrapolate to unseen, more heterogeneous datasets, underscoring the need for enhanced feature representation or transfer learning strategies to bridge the performance gap between historical and contemporary benchmarks.

In Table 8, the results obtained from the Davis dataset are presented and compared with the base approaches. The proposed method performs better in all metrics: the CI measure, R, MSE, and AUPR. Utilizing similar training samples in task prediction steps helps the approach achieve better results. Our approach improves the CI measure by 1.6% compared to the other comparable approaches. Additionally, it demonstrates a 7.6% improvement in AUPR compared to the base approaches. It should be noted that the boldfaced values indicate that the corresponding method outperforms the comparison methods.

thumbnail
Table 8. Comparing the proposed approach results with base approaches on the Davis, KIBA, and BindingDB Datasets.

https://doi.org/10.1371/journal.pone.0357278.t008

Table 7 also shows the achieved results of the proposed approach on the KIBA dataset. Also, it is compared with the results of the base approaches. In most measures, the proposed method provides the best results. It achieves a 5.9% improvement in AUPR over the second-best (DeepCDA). Our approach improves the CI measure by 1.7% compared to the other comparable approaches.

The results of our approach, along with its comparison to base approaches on the BindingDB dataset, are also presented in Table 8. Our approach yields the best results in this case across four measures. It shows a 5.9% improvement in AUPR. It should be noted that [51] utilizes smaller training and test sets. Their training set consists of 19312 samples, and the test set comprises 13289 samples. Our method and the other comparison methods in Table 7 use the same training set (with 101,134 samples) and test set (with 43,391 samples).

To statistically verify the proposed method, we have used the paired t-test. In this test, the null hypothesis () states that there is no difference between the proposed method and the comparison methods, and the alternative hypothesis () shows significant differences between them. As shown in the KIBA and BindingDB datasets, all reported p-values are below 0.05, indicating that we reject the null hypothesis and demonstrate a statistically significant difference between the proposed and comparable methods. It should be noted that to compute a reliable p-value, we require at least three values. This is the reason why we do not report the p-value for some approaches. Additionally, in all reported p-values, the MSE metric is not used because, in all metrics except MSE, a higher value is considered better.

5.3. Generalizability

In this section, the generalizability of the proposed approach is evaluated and applied to the BindingDB dataset. As mentioned in this dataset, four groups of proteins are entirely excluded from the training and test sets: nuclear estrogen receptors (516 samples), ion channels (8,101 samples), receptor tyrosine kinases (3,355 samples), and G-protein-coupled receptors (77,994 samples). As a result, an experiment is designed to evaluate how the proposed approach performs when a target protein is not seen during the training phase. The model is trained on a BindingDB dataset and then applied to nuclear estrogen receptors, ion channels, receptor tyrosine kinases, and G-protein-coupled receptors.

The results obtained from the other approaches are compared in Table 9. It shows that our approach in all datasets performs best in the CI, R, and AUPR measures compared to the other base approaches. Moreover, the proposed approach achieves an average improvement of 15.5% in AUPR. These results demonstrate that the proposed approach outperforms other methods in the classification task. Additionally, in the MSE metric, our approach outperforms DeepAffinity in three out of four datasets. For ER and Tyrosine kinase, all p-values are below 0.05, indicating that the proposed method outperforms DeepDTA and DeepCDA. Additionally, for the Ion channel and GPCR, all p-values are less than 0.08, suggesting that there's less than an 8% chance that the proposed method and the comparison methods are not different.

thumbnail
Table 9. Generalizability evaluation of our approach and base comparable approaches on BindingDB. In this case, the model is learned on the BindingDB dataset and evaluated on four other datasets.

https://doi.org/10.1371/journal.pone.0357278.t009

5.4. Target selectivity of ligands

To evaluate the proposed approach, we have done an experiment that assesses the target selectivity of ligands. In this experiment, we have a pool of ligands. Then, for each target, we compute the interaction affinity of the given target with a pool of ligand molecules to demonstrate how the proposed approach can predict the true order of interactions. We have chosen the Protein-tyrosine phosphatases (PTPs) family as protein targets. These protein families are involved in metabolic regulation and mitogenic signal transduction processes. The selected proteins in the PTP family are SHP1, PTPRC, PTPRE, PTPRA, and PTP1B. Also, we have used the following three ligands:

  1. 1) 2-(oxaloamino) benzoic acid with PubChem CID of 444764,
  2. 2) 3-(carboxyformamido) thiophene-2-carboxylic acid with PubChem CID of 44359299 and
  3. 3) 2-formamido-4,5,6,7-tetrahydrothieno [2,3-c]pyridine-3-carboxylic acid with PubChem CID of 90765696.

Recent studies show that all selected ligands are more attracted to the PTP1B protein. The values for the three ligands against PTP1B, respectively, are 4.63, 4.25, and 6.69. Also, the difference between value against PTP1B and the second closest protein (shown by ) respectively are 0.75, 0.7, and 2.47. The results obtained are presented in Table 10. Our approach predicts the affinity value in the same order as the ground truth one. Also, for the third compound, the predicted for our approach is closer to the true value compared with the other methods. To further evaluate the results, we have done molecular docking between the third ligand and the proteins that get the second-highest value against the third ligand in RA-PLA and DeepCDA approaches. In other words, molecular docking is done between two pairs: 1) molecular docking between the ligand with PubChem CID of 90765696 and the SHP1 protein, and 2) molecular docking between the ligand with PubChem CID of 90765696 and the PTPRA protein. To achieve this, we utilized AutoDock Vina. The docking scores reported in Table 11 represent calculated binding energies from AutoDock Vina (in kcal/mol), which serve as a computational estimate of binding strength. Lower (more negative) binding energies indicate stronger predicted binding, which is consistent with the ranking order predicted by RA-PLA and validated by AutoDock Vina's independent scoring. RA-PLA and AutoDock Vina predict the same order for the two pairs mentioned, while the predicted orders of DeepCDA and DeepDTA do not match those predicted by AutoDock Vina. Ligand efficiency is another metric reported, which expresses the binding energy per atom of the ligand to the protein. Also, the predicted ligand-protein complex for these two pairs is shown in Fig 3.

thumbnail
Table 10. The predicted value for compounds with PubChem CIDs of 444764 (C1), 44359299 (C2), and 90765696 (C3) against five proteins of the PTP family.

https://doi.org/10.1371/journal.pone.0357278.t010

thumbnail
Table 11. Molecular docking results between the PubChem CID of 90765696 drug and two proteins, SHP1 and PTPRA.

https://doi.org/10.1371/journal.pone.0357278.t011

thumbnail
Fig 3. The predicted ligand-protein complex.

(a) PubChem CID of 90765696 drug and SHP1 protein and (b) PubChem CID of 90765696 drug and PTPRA protein. AutoDock Vina does this prediction.

https://doi.org/10.1371/journal.pone.0357278.g003

5.5. Screening power

In this section, the screening power of the proposed method is assessed. The screening power test evaluates the capability of the scoring function in discriminating true binders from decoy compounds. This test is conducted on the PDBBind2016-core set (also known as the CASF-2016 dataset). It contains 285 high-quality protein-ligand complexes with 57 clusters of complexes such that in each cluster, five ligands are considered active targets, and the remaining ones (i.e., 280 ligands) are considered inactive. In the PDBBind2016-core set, two settings for screening power tests are designed: the forward screening power test and the reverse screening power test. In the forward screening test, a potential compound is identified for a specific target, and in the reverse screening test, a likely target is determined for a particular active compound. In this section, we conducted the forward screening test, and the results are reported in Table 12. To report the results, we utilized the Enrichment Factor (EF) as an evaluation measure, similar to state-of-the-art approaches. EF measure is defined as follows:

thumbnail
Table 12. Screening power results on PDBBind2016-core set (CASF-2016 dataset).

https://doi.org/10.1371/journal.pone.0357278.t012

(16)

where denotes the number of true binders among the x% top-ranked candidates and refers to the number of true binders among all candidates.

The introduced scoring functions are divided into two categories: classical scoring functions and machine learning (ML) based scoring functions. The enrichment factor for ML-based scoring functions is, on average, lower than that of classical approaches [57]. In Table 12, we have compared the proposed method with ChemScore@GOLD (Special implementation of [58,59]), Autodock Vina [60], ΔvinaRF20 [61], and Δ-AEScore [62]. Our approach gets results comparable to those of the classical methods. ΔvinaRF20 has obtained a better result compared to the other approaches. However, it is shown that in this approach, there is a 50% overlap between protein-ligand complexes of the training and test sets [57]. In the proposed approach, we use only the SMILES representation of the drug and the sequence of amino acids of the protein as input. In other words, we do not use the 3D complex of ligand-protein or the protein's 3D structure.

6. Discussion

Predicting protein-ligand interactions is crucial for new drug discovery and repurposing. A binding affinity measures how strong a protein-ligand interaction is. Biologists can reduce the time spent on wet laboratory experiments and accelerate drug discovery by computationally predicting the binding affinity between ligands and proteins [27,63–65]. This paper proposes an end-to-end deep learning-based approach for PLA prediction that utilizes the knowledge of hard samples of protein-ligand pairs in the affinity prediction step. One of the main advantages of the proposed method is that it leverages the observation that a single model has limited ability to fit an extensive feature space (in this case, the protein-ligand feature space). To address this issue, the proposed approach retrieves query-dependent hard training samples and incorporates them into inferring the label of the query sample. The proposed method first retrieves a hard sample of protein-ligand pairs for the input sample. Then, it is fed into the network to compute the features for the input protein-ligand pair and the retrieved pairs. Next, these descriptors are fed into the network to compute the adjacency matrix. Then, a graph is constructed using node descriptors and an adjacency matrix in which each node represents a protein-ligand pair. This graph is fed into a semi-supervised GCN to predict the label of the node corresponding to the query protein-ligand pair. Learning of the proposed network is done end-to-end. The proposed approach uses GCN not as a feature descriptor but as a node classification model.

The evaluation results of the proposed approach demonstrate its successful performance on the PDBbind, Davis, KIBA, and BindingDB datasets. One of the findings of the proposed approach is that it learns a model with more generalizability. The learned model on BindingDB is applied to four other datasets where their target proteins are not seen in the training phase (cold target setting). In other words, the mentioned target proteins are entirely excluded from BindingDB. The results confirm that our approach has more generalization ability than other approaches. Additionally, the proposed approach achieves better classification performance than regression.

One of the key motivations for this paper is that current approaches have limited generalizability due to relying on a single learned model to predict affinity values for all unseen protein-ligand pairs. We believe that the generalizability of this approach is limited. To consider this issue, we have designed a method that utilizes hard-retrieved samples in the task prediction step. To this end, a semi-supervised GCN is used. In this case, in the inference step, for an input sample, a graph is constructed in which most nodes (retrieved protein-ligand pairs) have labels, and the label of the input sample is unknown. This graph is fed into the semi-supervised GCN, and training is then performed. Finally, the label of the input samples is predicted by utilizing the knowledge of hard samples. The training step adjusts the parameter values correctly, and in the inference step, the learned model is fine-tuned for the input sample. This leads to the advantage of being partly independent of the dataset, as fine-tuning for each input sample is performed during the inference step. The results obtained from the BindingDB dataset confirm this statement.

One of the key reasons the proposed method leads to an improvement in performance is that, in the inference step, it not only utilizes the learned model from the training data but also fine-tunes the model with retrieved similar training samples for the test sample, resulting in a more accurate prediction. It means that in the inference step, we extract similar training samples for the test sample and construct a semi-supervised graph. Next, the model is fine-tuned using the constructed semi-supervised graph. Then, a prediction is made for the test sample. It is worth noting that we have a specific model for each test sample.

One of the other advantages of the proposed approach that helps it learn a more discriminative model is the retrieval step. In the retrieval step, the number of retrieved samples is different for each query input. If the query sample is placed near the decision boundary, the number of hard samples increases, and all these samples are incorporated into the decision-making process. In the training step, the number of retrieved samples is initially large and decreases as the iterations progress. The reason is that during the training process, the model becomes progressively more accurate and discriminative.

The proposed approach could help biologists deal with biological problems. In this case, we have designed a screening power experiment on the PDBBindv2016-core set (i.e., CASF 2016). This experiment evaluates the capability of the proposed method in discriminating true binders from decoy compounds. The results show that it achieves a comparable ranking even with the classical scoring function, which utilizes much of the information from the 3D complex of the protein-ligand complex.

One of the limitations of the proposed method is that the inference step for each test sample is slower than the other techniques. Because, for each test sample, the model should be fine-tuned using the retrieved training samples. However, we believe that achieving better performance is a higher priority than minimizing inference time in PLA prediction. Also, the maximum number of retrieved samples for each query should be limited due to memory limitations.

Acknowledgments

The authors thank the anonymous reviewers for their valuable suggestions.

References

  1. 1. Kianfar A, Razzaghi P, Asgari Z. Integrating convolutional layers and biformer network with forward-forward and backpropagation training. Sci Rep. 2025;15(1):7230. pmid:40021838
  2. 2. Dehghan A, Razzaghi P, Abbasi K, Gharaghani S. TripletMultiDTI: Multimodal representation learning in drug-target interaction prediction with triplet loss function. Expert Systems with Applications. 2023;232:120754.
  3. 3. Amiri M, Makkiabadi B. Knowledge-enhanced graph transformer with contrastive learning for drug-drug interaction prediction. Machine Learning with Applications. 2026;:100892.
  4. 4. Zhu Z, Zheng X, Qi G, Gong Y, Li Y, Mazur N, et al. Drug–target binding affinity prediction model based on multi-scale diffusion and interactive learning. Expert Systems with Applications. 2024;255:124647.
  5. 5. Song J, Yang Z, Song Z, Heng Y, Ge L, Zhang J, et al. RGGE-DTD: A Unified Model for Simultaneous Prediction of Drug-Target Interactions and Drug-Disease Associations in Drug Repositioning. Expert Systems with Applications. 2026;321:132319.
  6. 6. Das J, Tradigo G, Veltri P, Guzzi PH, Roy S. Data science in unveiling COVID-19 pathogenesis and diagnosis: evolutionary origin to drug repurposing. Briefings in Bioinformatics. 2021;22(2):855–72.
  7. 7. Zhou D, Xu Z, Li W, Xie X, Peng S. MultiDTI: drug-target interaction prediction based on multi-modal representation learning to bridge the gap between new chemical entities and known heterogeneous network. Bioinformatics. 2021;37(23):4485–92. pmid:34180970
  8. 8. Abbasi K, Razzaghi P, Gharizadeh A, Ghareyazi A, Dehnad A, Rabiee HR, et al. Computational drug design in the artificial intelligence era: A systematic review of molecular representations, generative architectures, and performance assessment. Pharmacol Rev. 2026;78(1):100095. pmid:41389438
  9. 9. Abbasi K, Razzaghi P, Poso A, Ghanbari-Ara S, Masoudi-Nejad A. Deep learning in drug target interaction prediction: current and future perspectives. Current Medicinal Chemistry. 2021;28(11):2100–13.
  10. 10. Sajadi SZ, Zare Chahooki MA, Gharaghani S, Abbasi K. AutoDTI++: deep unsupervised learning for DTI prediction by autoencoders. BMC Bioinformatics. 2021;22(1):204. pmid:33879050
  11. 11. Razzaghi P, Abbasi K, Shirazi M, Rashidi S. Multimodal brain tumor detection using multimodal deep transfer learning. Applied Soft Computing. 2022;129:109631.
  12. 12. Yu D, Liu H, Yao S. Drug–target interaction prediction based on improved heterogeneous graph representation learning and feature projection classification. Expert Systems with Applications. 2024;252:124289.
  13. 13. Schulman A, Rousu J, Aittokallio T, Tanoli Z. Attention-based approach to predict drug-target interactions across seven target superfamilies. Bioinformatics. 2024;40(8):btae496. pmid:39115379
  14. 14. Rush TS, Grant JA, Mosyak L, Nicholls A. A shape-based 3-D scaffold hopping method and its application to a bacterial protein− protein interaction. Journal of Medicinal Chemistry. 2005;48(5):1489–95.
  15. 15. Sutherland JJ, Higgs RE, Watson I, Vieth M. Chemical fragments as foundations for understanding target space and activity prediction. J Med Chem. 2008;51(9):2689–700. pmid:18386916
  16. 16. Salum LB, Andricopulo AD. Fragment-based QSAR strategies in drug design. Expert Opin Drug Discov. 2010;5:405–12.
  17. 17. Abbasi K, Razzaghi P, Poso A, Amanlou M, Ghasemi JB, Masoudi-Nejad A. DeepCDA: deep cross-domain compound-protein affinity prediction through LSTM and convolutional neural networks. Bioinformatics. 2020;36(17):4633–42.
  18. 18. Lv H, Dao F-Y, Zulfiqar H, Lin H. DeepIPs: comprehensive assessment and computational identification of phosphorylation sites of SARS-CoV-2 infection using a deep learning-based approach. Brief Bioinform. 2021;22(6):bbab244. pmid:34184738
  19. 19. Tian Q, et al. Predicting drug-target affinity based on recurrent neural networks and graph convolutional neural networks. Combinatorial Chemistry & High Throughput Screening. 2022;25(4):634–41.
  20. 20. Chen X, Song S, Ji J, Song S, Zhou X, Zhang H. An explainable machine learning-based scoring function using interpretable features and model explanation approaches for binding affinity prediction. Expert Systems with Applications. 2026;318:131919.
  21. 21. Öztürk H, Özgür A, Ozkirimli E. DeepDTA: deep drug-target binding affinity prediction. Bioinformatics. 2018;34(17):i821–9. pmid:30423097
  22. 22. Yuan W, Chen G, Chen CY-C. FusionDTA: attention-based feature polymerizer and knowledge distillation for drug-target binding affinity prediction. Brief Bioinform. 2022;23(1):bbab506. pmid:34929738
  23. 23. Song T, Wang G, Ding M, Rodriguez-Paton A, Wang X, Wang S. Network-Based Approaches for Drug Repositioning. Mol Inform. 2022;41(5):e2100200. pmid:34970871
  24. 24. Du B-X, Qin Y, Jiang Y-F, Xu Y, Yiu S-M, Yu H, et al. Compound-protein interaction prediction by deep learning: Databases, descriptors and models. Drug Discov Today. 2022;27(5):1350–66. pmid:35248748
  25. 25. Wang H, Wang J, Dong C, Lian Y, Liu D, Yan Z. A novel approach for drug-target interactions prediction based on multimodal deep autoencoder. Frontiers in Pharmacology. 2020;10:1592.
  26. 26. Lim J, Ryu S, Park K, Choe YJ, Ham J, Kim WY. Predicting Drug-Target Interaction Using a Novel Graph Neural Network with 3D Structure-Embedded Graph Representation. J Chem Inf Model. 2019;59(9):3981–8. pmid:31443612
  27. 27. Ru X, Ye X, Sakurai T, Zou Q. NerLTR-DTA: drug-target binding affinity prediction based on neighbor relationship and learning to rank. Bioinformatics. 2022;38(7):1964–71. pmid:35134828
  28. 28. Shim J, Hong Z-Y, Sohn I, Hwang C. Prediction of drug-target binding affinity using similarity-based convolutional neural network. Sci Rep. 2021;11(1):4416. pmid:33627791
  29. 29. Aleb N. Multilevel Attention Models for Drug Target Binding Affinity Prediction. Neural Process Lett. 2021;53(6):4659–76.
  30. 30. Nguyen T, Le H, Quinn TP, Nguyen T, Le TD, Venkatesh S. GraphDTA: predicting drug-target binding affinity with graph neural networks. Bioinformatics. 2021;37(8):1140–7. pmid:33119053
  31. 31. Son J, Kim D. Development of a graph convolutional neural network model for efficient prediction of protein-ligand binding affinities. PloS One. 2021;16(4):p.e0249404.
  32. 32. Liu X, Feng H, Wu J, Xia K. Dowker complex based machine learning (DCML) models for protein-ligand binding affinity prediction. PLoS Computational Biology. 2022;18(4):p.e1009943.
  33. 33. Nguyen DD, Wei G-W. AGL-Score: Algebraic Graph Learning Score for Protein-Ligand Binding Scoring, Ranking, Docking, and Screening. J Chem Inf Model. 2019;59(7):3291–304. pmid:31257871
  34. 34. Li Q, Zhang X, Wu L, Bo X, He S, Wang S. PLA-MoRe: A Protein-Ligand Binding Affinity Prediction Model via Comprehensive Molecular Representations. Journal of Chemical Information and Modeling. 2022.
  35. 35. Wang K, Zhou R, Li Y, Li M. DeepDTAF: a deep learning method to predict protein-ligand binding affinity. Briefings in Bioinformatics. 2021;22(5):bbab072.
  36. 36. Liu Z, Su M, Han L, Liu J, Yang Q, Li Y, et al. Forging the Basis for Developing Protein-Ligand Interaction Scoring Functions. Acc Chem Res. 2017;50(2):302–9. pmid:28182403
  37. 37. He T, Heidemeyer M, Ban F, Cherkasov A, Ester M. SimBoost: a read-across approach for predicting drug-target binding affinities using gradient boosting machines. J Cheminform. 2017;9(1):24. pmid:29086119
  38. 38. Davis MI, et al. Comprehensive analysis of kinase inhibitor selectivity. Nature Biotechnology. 2011;29(11):1046–51.
  39. 39. Tang J, Szwajda A, Shakyawar S, Xu T, Hintsanen P, Wennerberg K, et al. Making sense of large-scale kinase inhibitor bioactivity data sets: a comparative and integrative analysis. J Chem Inf Model. 2014;54(3):735–43. pmid:24521231
  40. 40. Mi J, Sun JH, Li J, Li C, Wan J. An artificial intelligence model for accurate drug-target affinity prediction in medicinal chemistry. Eur J Med Chem. 2026;311:118840. pmid:41950652
  41. 41. Liu T, Lin Y, Wen X, Jorissen RN, Gilson MK. BindingDB: a web-accessible database of experimentally determined protein-ligand binding affinities. Nucleic Acids Res. 2007;35(Database issue):D198-201. pmid:17145705
  42. 42. Liu Z, Su M, Han L, Liu J, Yang Q, Li Y, et al. Forging the Basis for Developing Protein-Ligand Interaction Scoring Functions. Acc Chem Res. 2017;50(2):302–9. pmid:28182403
  43. 43. Tang J, Szwajda A, Shakyawar S, Xu T, Hintsanen P, Wennerberg K, et al. Making sense of large-scale kinase inhibitor bioactivity data sets: a comparative and integrative analysis. J Chem Inf Model. 2014;54(3):735–43. pmid:24521231
  44. 44. Karimi M, Wu D, Wang Z, Shen Y. DeepAffinity: interpretable deep learning of compound-protein affinity through unified recurrent and convolutional neural networks. Bioinformatics. 2019;35(18):3329–38. pmid:30768156
  45. 45. Pahikkala T, et al. Toward more realistic drug–target interaction predictions. Briefings in Bioinformatics. 2015;16(2):325–37.
  46. 46. Ma W, Zhang S, Li Z, Jiang M, Wang S, Guo N, et al. Predicting Drug-Target Affinity by Learning Protein Knowledge From Biological Networks. IEEE J Biomed Health Inform. 2023;27(4):2128–37. pmid:37018115
  47. 47. Liu Y, Xing L, Zhang L, Cai H, Guo M. GEFormerDTA: drug target affinity prediction based on transformer graph for early fusion. Sci Rep. 2024;14(1):7416. pmid:38548825
  48. 48. Mazzone E, Moreau Y, Fariselli P, Raimondi D. Nonlinear data fusion over Entity-Relation graphs for Drug-Target Interaction prediction. Bioinformatics. 2023;39(6):btad348. pmid:37255310
  49. 49. Walsh I, Fishman D, Garcia-Gasulla D, Titma T, Pollastri G, ELIXIR Machine Learning Focus Group, et al. DOME: recommendations for supervised machine learning validation in biology. Nat Methods. 2021;18(10):1122–7. pmid:34316068
  50. 50. Cang Z, Mu L, Wei GW. Representability of algebraic topology for biomolecules in machine learning based scoring and virtual screening. PLoS Computational Biology. 2018;14(1).
  51. 51. Kang H, Goo S, Lee H, Chae J-W, Yun H-Y, Jung S. Fine-tuning of BERT Model to Accurately Predict Drug-Target Interactions. Pharmaceutics. 2022;14(8):1710. pmid:36015336
  52. 52. D’Souza S, Prema KV, Balaji S, Shah R. Deep Learning-Based Modeling of Drug-Target Interaction Prediction Incorporating Binding Site Information of Proteins. Interdiscip Sci. 2023;15(2):306–15. pmid:36967455
  53. 53. Zhang L, Zeng W, Chen J, Chen J, Li K. GDilatedDTA: Graph dilation convolution strategy for drug target binding affinity prediction. Biomedical Signal Processing and Control. 2024;92:106110.
  54. 54. Shah PM, Zhu H, Lu Z, Wang K, Tang J, Li M. DeepDTAGen: a multitask deep learning framework for drug-target affinity prediction and target-aware drugs generation. Nat Commun. 2025;16(1):5021. pmid:40447614
  55. 55. Li Z, Zeng Y, Jiang M, Wei B. Deep Drug-Target Binding Affinity Prediction Base on Multiple Feature Extraction and Fusion. ACS Omega. 2025;10(2):2020–32. pmid:39866608
  56. 56. Ni Z, Wei B, Zeng Y. MGF-DTA: A Multi-Granularity Fusion Model for Drug-Target Binding Affinity Prediction. Int J Mol Sci. 2026;27(2):947. pmid:41596595
  57. 57. Su M, Yang Q, Du Y, Feng G, Liu Z, Li Y, et al. Comparative Assessment of Scoring Functions: The CASF-2016 Update. J Chem Inf Model. 2019;59(2):895–913. pmid:30481020
  58. 58. Baxter CA, Murray CW, Clark DE, Westhead DR, Eldridge MD. Flexible docking using Tabu search and an empirical estimate of binding affinity. Proteins. 1998;33(3):367–82. pmid:9829696
  59. 59. Eldridge MD, Murray CW, Auton TR, Paolini GV, Mee RP. Empirical scoring functions: I. The development of a fast empirical scoring function to estimate the binding affinity of ligands in receptor complexes. Journal of computer-aided molecular design. 1997;11:425–45.
  60. 60. Trott O, Olson AJ. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J Comput Chem. 2010;31(2):455–61. pmid:19499576
  61. 61. Wang C, Zhang Y. Improving scoring-docking-screening powers of protein-ligand scoring functions using random forest. J Comput Chem. 2017;38(3):169–77. pmid:27859414
  62. 62. Meli R, Anighoro A, Bodkin MJ, Morris GM, Biggin PC. Learning protein-ligand binding affinity with atomic environment vectors. J Cheminform. 2021;13(1):59. pmid:34391475
  63. 63. Pan F, Yin C, Liu S-Q, Huang T, Bian Z, Yuen PC. BindingSiteDTI: differential-scale binding site modelling for drug-target interaction prediction. Bioinformatics. 2024;40(5):btae308. pmid:38730554
  64. 64. Tang X, Lei X, Zhang Y. Prediction of Drug-Target Affinity Using Attention Neural Network. Int J Mol Sci. 2024;25(10):5126. pmid:38791165
  65. 65. Zhong Y, Li G, Yang J, Zheng H, Yu Y, Zhang J, et al. Learning motif-based graphs for drug–drug interaction prediction via local–global self-attention. Nat Mach Intell. 2024;6(9):1094–105.