Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

T-pGNN4DTI: Towards better drug-target interactions prediction using Global Self-attentive Pooled Graph Convolutional Networks and protein pre-training Models

  • Yanmei Lin ,

    Roles Conceptualization, Formal analysis, Methodology, Writing – original draft

    ☯ Yanmei Lin and Boqi Yang are co-first authors and contributed equally to this work.

    Affiliation College of Big Data and Software Engineering, Zhejiang Wanli University, Ningbo, China

  • Boqi Yang ,

    Roles Data curation, Methodology, Software, Writing – review & editing

    ☯ Yanmei Lin and Boqi Yang are co-first authors and contributed equally to this work.

    Affiliation College of Arts and Sciences, Emory University, Atlanta, United States of America

  • Jianping Liao,

    Roles Conceptualization, Formal analysis

    Affiliation Guangxi Key Lab of Human-machine Interaction and Intelligent Decision, Nanning Normal University, Nanning, China

  • Chenjie Du,

    Roles Supervision, Writing – review & editing

    Affiliation College of Big Data and Software Engineering, Zhejiang Wanli University, Ningbo, China

  • Hongguo Cai,

    Roles Formal analysis, Writing – review & editing

    Affiliation Guangxi Key Lab of Human-machine Interaction and Intelligent Decision, Nanning Normal University, Nanning, China

  • Yijia Wu ,

    Roles Methodology, Software, Supervision, Validation, Writing – original draft

    wuyijia06@163.com (YW); PengYZ@zwu.edu.cn (YP)

    Affiliation Guangdong University of Technology, Zhaoqing, China

  • Yuzhong Peng

    Roles Project administration, Supervision, Writing – review & editing

    wuyijia06@163.com (YW); PengYZ@zwu.edu.cn (YP)

    Affiliation College of Big Data and Software Engineering, Zhejiang Wanli University, Ningbo, China

Abstract

Identification of drug-target interactions (DTI) is an important and challenging task in drug discovery and development. Traditional methods generally require biological experiments, which are costly and time-consuming. Machine learning-based methods can rapidly predict DTI using only computer algorithmic models, allowing researchers to validate only the most promising interactions through biochemical experiments. This holds promise for effectively addressing the current challenges of lengthy development cycles and high costs in new drug development. However, it is difficult for the existing DTI prediction methods to learn complete and effective feature information from the compound and protein. Therefore, this work proposes a DTI prediction method based on the global self-attentive pooled graph neural network and protein pretraining model, called T-pGNN4DTI. On the one hand, T-pGNN4DTI uses a global self-attention pooled graph neural network to learn more meaningful features of the drug molecule by paying more attention to the information features of certain important atomic nodes of the molecular structure and ignoring some weakly relevant node information features. On the other hand, T-pGNN4DTI uses a pre-trained Transformer-based model to capture the semantic relationships of contexts in long sequences of proteins, which can learn more complete feature information. The results of comparing experiments on three benchmark datasets show that the performance of the proposed T-pGNN4DTI model is better than that of the existing DTI prediction methods, effectively improving the DTI prediction. It provides a new way of thinking to help solve the DTI-related problems.

Introduction

Most drug compounds achieve their therapeutic effects by interacting with specific target molecules (proteins) within the human body to regulate their functions [1]. Therefore, identification of the DTI is one of the most important tasks in the drug discovery process, yet this process using conventional biological experiment-based methods is extremely time-consuming and costly [24]. In recent years, advances in artificial intelligence (AI) have enabled machine learning methods to successfully address several complex challenges in drug research and discovery [510], driving significant progress in related fields. Consequently, an increasing number of researchers are exploring the application of machine learning methods to address a broader range of complex problems in drug discovery, including DTI prediction. The emergence of AI-based drug-target interaction prediction methods offers an opportunity to alleviate these challenges [3,4,8]. By predicting outcomes, researchers can better explore and understand the mechanisms of drug action, revealing how drugs modulate biological processes to produce therapeutic effects, thereby providing valuable guidance for the design and discovery of new drugs. Accurate and efficient DTI prediction can effectively speed up the virtual screening process of potential compounds [11]. They can minimize unnecessary biological experiments by narrowing the search scope of potential compounds, thus accelerating the development of new drugs [12].

Traditional machine learning methods for DTI prediction can be classified into two categories: similarity-based methods and feature-based methods. Similarity-based methods include Matrix Factorization (MF)-based methods and kernel-based methods. For example, Zheng et al. [12] proposed a Collaborative Matrix Factorization (CMF) based method for DTI prediction; Cichonska et al. [13] proposed a Kronecker product based on compound and protein kernels for DTI prediction. He et al. [14] proposed a classical feature-based DTI prediction method called SimBoost, which defines three types of features for drug, target, and drug-target pairs, respectively, each of which contains multiple hand-created features. These traditional machine learning-based methods, despite achieving good advances in DTI prediction, are still far from the expectations of engineering applications. Numerous studies indicate that deep learning methods outperform traditional machine learning methods in many domains, including drug-target interaction prediction, and hold greater potential for future development [11].

Deep learning-based DTI prediction methods generally include two steps: first, using deep neural networks to learn the features of compounds and proteins from the input raw data; then, the obtained features are used to make classification predictions after completing the trained model. Currently, scholars have developed Convolutional Neural Network (CNN)-based models for feature extraction of compounds and proteins [11,15,16]. However, their difficulty in capturing long-range dependencies results in significant limitations in molecular feature extraction and prevents the representation of potential interactions between atoms at long distances based on the original molecular sequence. To learn more comprehensive drug features, Graph Neural Networks (GNNs) and their variants have also been proposed for extracting features from the 2D structure of compounds [1719]. For example, Graph Attention Networks (GAT) [20] and Gated Graph Sequence Neural Networks [21] have been used for DTI prediction tasks. Hu et al. [18] proposed a graph neural network-based DTI prediction model, iGRLDTI, which employs an enhanced graph representation learning method to more effectively capture discriminative representations of drugs and targets in the latent feature space, aiming to address the issue of over-smoothing during the simulation process. In recent years, researchers have also employed combinations of different neural networks or attention mechanisms to enhance the accuracy of DTI predictions [22,23]. For example, DeepCDA [24] employs a combination of Long Short-Term Memory (LSTM) and CNN to predict DTI, which represents compounds as text sequences composed of “words” and encodes proteins using word embedding techniques, enabling the LSTM to extract effective features of compounds and proteins from massive unlabeled corpora. DrugBAN [25] utilizes graph convolutional networks (GCNs) and one-dimensional CNNs to extract substructure features from drug molecular graphs and protein sequences, respectively. It then employs a bilinear attention network module to explicitly learn the local interaction relationships between drug-target pairs. AttentionDTA [26] incorporates an attention module after the drug and protein feature extraction stages to emphasize specific features, thereby optimizing DTI prediction outcomes. The application of pre-trained language models (LMs) has become a powerful tool across multiple research domains. BERT (Bidirectional Encoder Representations from Transformers) [27] sparked a paradigm shift in natural language processing tasks, with its influence extending beyond this field. Some pre-trained models for proteins and chemical compounds have been applied in Drug-Drug Interaction (DDI) prediction [28] and DTI prediction studies [2932], where language models are utilized to generate embedding vectors.

Overall, despite advances in deep learning-based DTI prediction methods, existing approaches still face limitations in feature representation learning and outcome prediction. For instance, when using GCN to learn features of drug compounds, the network treats all atomic nodes as equivalent and assigns them equal weight. This results in certain critical nodes receiving insufficient attention, thereby hindering the identification of key atomic nodes within compound molecular structures. In the process of protein feature extraction, models based on CNN that employ one-hot encoding to extract protein features also exhibit limitations. The primary reason lies in the fact that each protein sequence requires preprocessing before one-hot encoding: longer sequences are truncated, while shorter sequences necessitate padding. This results in the features learned by the subsequent CNN being incomplete. The existing pretrained models generate independent embeddings that disregard neighborhood information. Furthermore, the previous language model-based DTI prediction studies focused solely on comparing different language model variants, lacking comprehensive comparisons with other methods [33].

To tackle the shortcomings of current DTI prediction methods, we studied optimizing feature representation and learning, then proposed a DTI prediction method based on a global self-attention graph convolutional network and a protein pre-training model (T-pGNN4DTI) to empower drug-target interactions prediction.

Materials and methods

Benchmark datasets

The dataset employed in this study follows TransformerCPI [34], drawing from three publicly available benchmark datasets widely recognized in the DTA prediction field in recent years [19,35], including the BindingDB, Human, and C.elegans datasets. Positive samples for the Human and C.elegans datasets were derived from positive drug-protein interaction (CPI) pairs in DrugBank 4.1 and Matador [34], with high-confidence negative samples generated through the negative CPI screening framework [35]. A more detailed introduction to the three benchmark datasets is given below.

BindingDB is a publicly accessible large-scale database measuring binding affinities, primarily focusing on interactions between proteins considered drug targets and drug-like small molecules [36]. The BindingDB dataset used in this work was obtained from TransformerCPI [34], containing 39747 positive interactions (involving 1696 protein targets and 53253 small molecules) and 31218 negative samples. The C.elegans dataset contains extensive information on gene, protein, and compound interactions in the nematode Caenorhabditis elegans, including 4000 positive interactions between 1434 unique compounds and 2,504 unique proteins specific to the C.elegans species. The Human dataset contains 852 drugs and 1052 proteins, covering 3369 interaction pairs. Table 1 summarizes the statistics of compounds, proteins, and the number of positive and negative samples in each dataset.

Overview of the model

The framework of T-pGNN4DTI, shown in Fig 1, consists of four main parts: the data preprocessing module, the feature extraction module, the feature fusion module, and the output module. (1) The data preprocessing module processes the compound data represented by SMILES strings and constructs the corresponding molecular maps of the compounds via the RDKit toolkit. (2) The feature extraction module mainly consists of a global self-attention pooled graph convolutional neural network for learning compound features from molecular maps and a pre-trained model for learning target features from protein sequences. (3) The feature fusion module fuses the features output from the graph convolutional network with the features output from the protein pre-training model for further learning. (4) The output module, consisting of a fully connected layer and a dropout layer, outputs the predicted values of compound-protein pairs.

thumbnail
Fig 1. The framework of the proposed T-pGNN4DTI model.

T-pGNN4DTI primarily comprises a data preprocessing module (for preprocessing drug data and target protein data, respectively), a feature extraction module (for extracting features from drug data and protein data, respectively), a feature fusion module, and an output module.

https://doi.org/10.1371/journal.pone.0352250.g001

The technical details of the main modules are as follows:

Data preprocessing module

Based on the SMILES of each compound, the tool RDKit was used to construct the corresponding molecular graph reflecting its structural information and inter-atomic interactions, as shown in Fig 2. In the structured molecular graph data, the information of each node can be mapped into a multidimensional feature vector, which mainly describes five pieces of information: the atom symbols, the number of neighboring hydrogens, the number of neighboring atoms, the implied value of the atom, and whether the atom is in an aromatic structure.

thumbnail
Fig 2. The process diagram of drug SMILES transformation.

T-pGNN4DTI utilizes the RDKit to convert drug SMILES into molecular graph data, which contains molecular structural relationships and molecular-related attribute information.

https://doi.org/10.1371/journal.pone.0352250.g002

T-pGNN4DTI uses a Transformer-based protein pre-training model to learn protein features without processing raw protein data in the data preprocessing module.

Feature extraction module

Feature of the protein extraction module.

With the development of deep learning technology, Transformer-based large-scale pre-trained models (PTMs) have achieved significant success in recent years in various fields, including natural language processing and computer vision. PTMs can efficiently extract knowledge from massive amounts of labeled and unlabeled data, storing this knowledge within a vast parameter system. After fine-tuning for specific downstream tasks, they can support a wide range of downstream applications. On the other hand, protein sequences can also be viewed as a special type of text-based language data, providing favorable conditions for applying language pre-training models to learn protein sequence features.

Inspired by ESM [37], we use a transformer-based protein pre-training model, PTR, for learning protein features in this work. The model is a bidirectional Masked language model, enabling it to observe information (context) near a specific residue or motif and thereby predict that residue or motif. It comprises an encoder module and a decoder module, which are both based on multi-head attention mechanisms and feedforward neural networks, as illustrated in Fig 3. The work process of PTR can be divided into two stages: the pre-training stage and the application stage. The former stage provides the latter stage with a large number of model parameters and contextual relationships within sequences.

thumbnail
Fig 3. The transformer-based protein pre-training model comprises an encoder block and a decoder module block.

Both encoders and decoders process inputs through a sequence of modules alternating between multi-head self-attention layers and feedforward layers.

https://doi.org/10.1371/journal.pone.0352250.g003

In the pre-training phase, the first step is to divide the original protein sequence into several subsequences in terms of characters, starting with the token [CLS] and ending with the token [SEP], . In the second step, each input vector is obtained by summing the character vectors (Token Embeddings, TE) and character position vectors (Position Embeddings, PE). Where the character vectors are computed using the Word2Vec model [38]. The feed-forward neural network of Word2Vec is used to convert each ID , dividing the input sequence into an embedding vector . The position vector is denoted by PE and is computed using the following equation:

(1)(2)

where pos denotes the position of the character in the sequence. d denotes the dimension of the PE, 2i denotes the even dimension, and 2i + 1 denotes the odd dimension (). In the third step, some of the TEs are randomly masked before the vectors are input to the encoder block, and the masked vectors are labeled with MASK. Then the masked characters are predicted by combining the context of the sequence:

(3)

where M is the set of masked characters, represents the number of characters in the protein sequence. For each input sequence, a certain proportion of characters is masked, and predictions are then made for the masked portions based on contextual information within the sequence. After extensive training, the model learns to recognize dependencies between masked and unmasked segments within the sequence, thereby enabling more accurate characterization of proteins. In this way, the relationship between contexts that are far apart in the protein sequence can be learned effectively, and features of the protein sequence can be better extracted. Finally, the embedding vector of the protein sequence is used as input to the PTR model. In the PTR model, IV is first fed into the multi-head self-attention layer of the encoder. Within each head of the multi-head self-attention layer, we project the protein features onto query, key, and value vector spaces via projection matrices. The output of each attention head is the result of scaled dot-product attention, as follows,

(4)

where is the self-attention score matrix, which is an attention weight matrix that follows a probability distribution; is the scaling factor (typically set to the dimension of the key vector); Q, K, and V represent the query matrix, key matrix, and value matrix, respectively. They are obtained by multiplying the input features by their corresponding weight matrices. The encoder runs the scaled dot-product attention mechanism demonstrated by Eq. (4) M (i.e., the number of heads) times with M different queries in parallel and concatenates their outputs. In this way, the self-attention mechanism explicitly constructs pairwise interactions () between all positions in the sequence, enabling the model to represent interactions between residues directly. Additionally, the PTR model adopted a pre-training strategy similar to AlphaFold3.The pre-training corpus is downloaded from RCSB PDB (https://www.rcsb.org, including more than 250000 structures from the PDB archive and 1060000 Computed Structure Models), as well as the protein Swiss-Prot database (https://www.uniprot.org, including about 570000 structures).

In the application phase, similar to the pre-training phase, proteins are first segmented into character-based subsequences. Then, these subsequences are encoded and fed into the model. The huge number of parameters and the contextual correlations in the sequences learned in the pre-training phase are applied when protein features are extracted. Finally, we utilize the feature matrix output from the last layer of the PTR model as the protein representation.

Drug feature extraction

In recent years, Junhyun Lee et al. [39] proposed Self-Attention Graph Pooling (SAGP), a layer different from the traditional pooling technique in graph neural networks. This method processes input image data by evaluating the relational importance of each node to its neighbors and aggregating based on importance weights, achieving superior performance across numerous image classification tasks. Benefiting from its dynamic weight allocation and hierarchical feature extraction techniques, it demonstrates significant advantages over traditional graph neural networks in noise robustness, structural fault tolerance, and data missing compensation. Inspired by it, to more comprehensively capture the intricate interactions among atoms, between atoms and motifs, and between atoms and motif-functional groups, we propose a Global Self-Attention Pooling-based Graph Convolutional Network (GSAP-GCN). This approach enhances the learning of compound molecular features—including node features and topological structure features—to improve drug-target interaction prediction. The GSAP-GCN framework is shown in Fig 4.

thumbnail
Fig 4. The semantic diagram of the proposed graph convolutional neural network with global self-attention pooling.

It describes the pipeline of the proposed graph convolutional neural network.

https://doi.org/10.1371/journal.pone.0352250.g004

Given a molecular graph of a chemical compound, G=(V, E). V is a set of nodes (i.e., atoms) described by the feature matrix X (representing the attributes or features of nodes within the molecular graph). Each node has its own feature vector represented by a C-dimensional vector (C denotes the number of features of a given node). E is the set of edges (i.e., bonds) described by the adjacency matrix A (representing the connections between nodes within the molecular graph). The main process of GSAP-GCN to learn the features of this compound is described as follows:

  1. (1) Construct the Node Feature Matrix of the molecular graph of compounds and the adjacency matrix , which are then inputted to the first graph convolution layer to compute the output node primary feature vectors denotes the node feature dimension). The computation rule of the graph convolution layer is defined as in Eq (5): (5)
    where denotes the feature matrix in layer l; denotes the weight matrix of layer l; is the augmented adjacency matrix with self-connectivity, and is the degree matrix of the node. A new feature matrix is first obtained by linearly transforming the multiplication by and . Then the weighted sum of the nodes and their neighbors is computed by normalizing the augmented adjacency matrix (i.e., . Finally, a feature matrix in the l + 1 layer H(l+1) is obtained by performing another nonlinear transformation through an activation function .
  2. (2) The node feature vectors output from the first graph convolutional layer are input to a second graph convolutional layer for similar computational processing before being input to a third graph convolutional layer, and so on. In particular, a node feature concatenation layer is attached after the last graph convolutional layer, and the output features of each graph convolutional layer are also simultaneously delivered to this concatenation layer.
  3. (3) The concatenation results from the node feature concatenation layer are input into the pooling layer with the self-attention mechanism. The global self-attention pooling layer is capable of filtering nodes, mainly including scoring the importance of nodes and discarding some node information through the masking mechanism. Fig 5 shows the processing of nodes by the global self-attention pooling layer. For ease of description, we denote the input feature matrix of the molecule in the global self-attention pooling layer as . First, to enable simultaneous consideration of node features and graph topological features during molecular information aggregation, our global self-attention pooling layer employs graph convolutions to compute the self-attention scores matrix , as demonstrated in Eq (6). (6)
    where is a learnable convolutional weight matrix; is the activation function (the Leaky ReLU function in this work). Alternatively, the softmax function can be used to normalize Z, though this comes at the cost of computational efficiency. This calculation method of the self-attention scores matrix differs from the aforementioned PTR network, which calculates self-attention scores using similarity measures that consider only node features, as shown in Eq (4).
    Then, a node pruning pooling method based on the self-attention scores ranking that prioritizes discarding relatively unimportant nodes while retaining only those deemed significant. Specifically, the number of previous [kN] nodes is retained based on the order of the self-attention scores. represents an adjustable SAG pooling ratio, which is a crucial hyperparameter determining the final model’s prediction performance. The formula for discarding nodes is defined as in Eqs (7) and (8): (7) (8)
    where, is the function that ranks by U and selects the index value of the former R nodes; idx is the index value; is the attention mask of the feature. Attention mask calculation output process, as shown in Eq (9): (9)
    where is the feature matrix indexed by node; ⊙ is the broadcasted element-wise product; is the row-wise and column-wise indexed adjacency matrix; and are the new feature matrix and the corresponding adjacency matrix computed by the attention mask, respectively. In this way, GSAP-GCN can simplify the graph structure while preserving its core structural features and key information, thereby reducing the graph’s size and enhancing computational efficiency.
  4. (4) The retained nodes are then aggregated through the Readout layer to obtain a representation of the graph, which is the feature of the compound. The information of the nodes retained by the global self-attentive pooling layer will be aggregated through the Readout layer. The aggregation is calculated as in Eq (10): (10)
    where N is the total number of nodes; denotes the feature vector of the th node and denotes the concatenation operation.
  5. (5) Finally, the aggregated representation vector of the obtained graph is fed into the feed-forward neural network with Dropout operation for further learning and optimization, thereby obtaining more effective drug features .
thumbnail
Fig 5. The working mechanism of the global self-attention pooling layer.

It illustrates the processing of molecular graph information within the global self-attention pooling layer, where Z denotes the attention score.

https://doi.org/10.1371/journal.pone.0352250.g005

Feature fusion and output module

In this module, we fuse the compound feature vectors learned from the graph convolutional network with the vectors of protein features output from the pre-trained model for interaction prediction. Specifically, for each compound-protein pair, the drug feature vector learned from the global self-attentive graph neural network and the protein feature vector learned from the PTR network are fused into the feature vector . Finally, the model input to the feed-forward neural network with Dropout operation and Softmax output layer, which is used for drug-target interaction prediction, as shown in Eq (11):

(11)

where and are the weight matrix and bias vector, respectively. out is the predicted label of the final output compound-protein pairs.

Results and discussion

Experimental Setup

Evaluation metrics.

DTI prediction is a classification task, so we employ metrics such as AUC (Area Under the Receiver Operating Characteristic curve), PRC (Area Under the Precision-Recall Curve), Precision, and Recall to evaluate model performance. Precision and Recall are described by the following formulas:

(12)(13)

where TP, FP and FN are represent the number of true positive, false positive and false negative samples, respectively.

Baseline models

In this work, the proposed model is compared with state-of-the-art methods, including SAG-DTA [40], TransformerCPI [34], GraphDTA [17],SAG(including SAGglobal and SAGhierarchical) [40], CPGL [41], AMMVF-DTI [42], MCL-DTI [43], TopoPharmDTI [44], IMAEN [45], MGNDTI [46], CF-DTI [47], and FusionDTI [48] on different datasets.

To verify the advantages of the global self-attention pooling technique used in this work, a T-pGNN4DTI variant model (T-hGNN4DTI) replacing the global self-attention pooling technique with the hierarchical self-attention pooling technique [39] is constructed experimentally to incorporate the SAGhierarchical method. SAGhierarchical can be divided into three modules, each containing a graph convolution layer and a pooling layer with a self-attention mechanism. The node features obtained from each module are aggregated in the Readout layer. The aggregated node features are summed and then passed through a fully connected layer to obtain a representation of the graph, which is the features of the compound.

In addition, we conducted some t-tests for analyzing the significance of the comparative models’ performance to scientifically determine model superiority.

Hyperparameter setting

In this work, we followed the dataset partitioning of SAG-DTA [40] that, the BindingDB dataset is partitioned into 80% training set, 10% testing set and 10% validation set, while the other datasets are partitioned into 80% training set and 20% testing set. Hyperparameters used in T-pGNN4DTI and its variant are shown in Table 2, while hyperparameters for state-of-the-art methods remain consistent with the settings in the original literature.

thumbnail
Table 2. Hyperparameters used in T-pGNN4DTI and its variant.

https://doi.org/10.1371/journal.pone.0352250.t002

Experimental results

In this section, the experimental results for our T-pGNN4DTI and its variant models (T-hGNN4DTI), GCN, FusionDTI, SAGglobal, and SAGhierarchical are averages obtained from 10 independent runs with different random seeds, respectively. Except for the comparative results of GCN, FusionDTI, SAGglobal, and SAGhierarchical, which were obtained through our repeated experiments, the comparison results for other compared models are cited from the corresponding references.

Tables 3–5 show the performance with P-values of our T-pGNN4DTI model against state-of-the-art methods on the C.elegans, Human, and BindingDB datasets.

thumbnail
Table 3. Comparison results on the C.elegans dataset.

https://doi.org/10.1371/journal.pone.0352250.t003

thumbnail
Table 4. Comparison results on the Human dataset.

https://doi.org/10.1371/journal.pone.0352250.t004

thumbnail
Table 5. Comparison results on the BindingDB dataset.

https://doi.org/10.1371/journal.pone.0352250.t005

Comparisons results

On the C.elegans dataset, as shown in Table 3, T-pGNN4DTI outperforms all state-of-the-art models across all performance metrics. It achieves over 0.20% to 2.46% higher AUC values, 0.60% to 6.40% higher PRC values, 1.43% to 8.03% higher Precision values, and 0.91% to 9.10% higher Recall values than the state-of-the-art models, respectively. Notably, the variant T-hGNN4DTI achieved second place, its all performance metrics only slightly lower than T-pGNN4DTI, while also outperforming all state-of-the-art models across all performance metrics on the C.elegans dataset. This indicates that the DTI prediction method using the pooled graph neural network to learn drug features and using the protein pretraining model to learn target features may be a better way than existing methods.

On the Human dataset, as shown in Table 4, T-pGNN4DTI achieves the best AUC, PRC, and Precision values in comparison with state-of-the-art methods. It achieves over 0.30% to 3.77% higher AUC values, 0.10% to 10.99% higher PRC values, and 0.21% to 9.98% higher Precision values than the state-of-the-art models, respectively. For Recall, MCL-DTI achieved the best value (0.961), which is 1.56% higher than that of our T-pGNN4DTI. Overall, T-pGNN4DTI outperforms state-of-the-art methods in terms of AUC, PRC, and Precision, while only slightly underperforming MCL-DTI in terms of Recall. Notably, the variant T-hGNN4DTI achieved second place in terms of AUC, PRC and Precision, while ranking third in terms of Recall on the Human dataset.

On the BindingDB dataset, as shown in Table 5, our T-pGNN4DTI outperforms all state-of-the-art models in terms of AUC, PRC, and Precision. It achieves over 0.93% to 61.19% higher AUC values, 0.52% to 78.82% higher PRC values, and 1.89% to 9.43% higher Precision values than the state-of-the-art models, respectively. For Recall, SAGglobal achieved the best value (0.942), which is 2.06% higher than that of our T-pGNN4DTI. Overall, T-pGNN4DTI outperforms state-of-the-art methods in terms of AUC, PRC, and Precision, while only slightly underperforming SAGglobal in terms of Recall. It achieves a big improvement in terms of AUC and PRC. Notably, the variant T-hGNN4DTI achieved second place in terms of AUC, PRC and Recall on the BindingDB dataset.

Case and Interpretability for T-pGNN4DTI

This study employs GSAP-GCN for feature representation learning of compounds, achieving not only excellent DTI prediction performance but also enhancing model interpretability. Specifically, hierarchical pooling dynamically adjusts the contributions of node neighborhoods via attention mechanisms to obtain concise molecular feature vectors that are rich in crucial molecular graph information. By observing the attention coefficients during molecular graph learning, we can measure the relevance of specific node embeddings and neighborhood embeddings to the prediction task. To illustrate this mechanism, we demonstrate and analyze the model’s determination of drug-target interactions between Ritonavir and HIV-1 protease. As shown in Figs 6 and 7

thumbnail
Fig 6. Visualization of Ritonavir feature learning and attention in GSAP-GCN.

This picture demonstrates the complete process of GNN learning the molecular features of ritonavir, divided into five subfigures: (A) GSAP-GCN processing flowchart for ritonavir molecules, (B) Pooling process schematic diagram, (C) Attention distribution scatter plot subfigure, (D) Full attention visualization subgraph, and (E) Bar chart for high correlation between GSAP-GCN attention and binding energy.

https://doi.org/10.1371/journal.pone.0352250.g006

thumbnail
Fig 7. Visualization of the T-pGNN4DTI model learning to model the interaction between Ritonavir and HIV-1 protease.

Subfigure (A) describes HIV-1 Protease Structure & Binding Pockets (Ritonavir Binding Sites) in the medical field. Subfigure (B) describes Ritonavir Pharmacophores captured by the attention of GSAP-GCN.

https://doi.org/10.1371/journal.pone.0352250.g007

In Fig 6, Subfigure (A) describes GSAP-GCN processing flowchart for ritonavir molecules, which includes three steps: (1) Molecular Input, (2) Node Attention Calculation, and (3) Attention Pooling – Motif-Level Aggregation. Step 1 displays the original structure of ritonavir, with key chemical motifs—including Thiazole, Amide, Hydroxyl, Phenyl A, and Phenyl B—highlighted by dashed lines in different colors. Step 2 displays the atomic-level attention scores output by the GNN layer. Colors range from green (low attention) to red (high attention); node size is proportional to attention scores; edge thickness indicates message-passing weights. Step 3 displays aggregating atomic attention scores into motif importance scores. The thiazole ring achieves the highest score (0.92), indicating its status as a core functional group. This is consistent with pharmaceutical facts in the medical field. It indicates that our model can accurately learn the key motif through which the Ritonavir molecule interacts with HIV-1 protease.

Subfigure (B), a schematic diagram of the pooling process, illustrates the workflow from atom → motif → molecular embedding → bioactivity prediction.

Subfigure (C), an attention distribution scatter plot, displays attention distributions grouped by motif, showing the mean ().

Subfigure (D) describes an overall visualization of GSAP-GCN attention, displaying (1) attention scores for atomic nodes, where redder colors indicate higher attention and yellow halos denote key atoms (≥ 0.8); (2) attention scores for chemical bonds, where thickness represents learned message-passing weights; and (3) attention scores for motifs, where dashed ellipses circle important regions and aggregate scores are annotated. Key insights from observing the subfigure (B) include: (1) The thiazole ring (attention score 0.92) is critical for binding to the HIV protease active site; (2) The hydroxyl group (attention score 0.88) serves as a hydrogen bond donor essential for binding; (3) The amide bond (attention score 0.82) plays an important role in binding. This is consistent with pharmaceutical facts in the medical field.

Subfigure (E), the high correlation bar chart, shows the high correlation between GSAP-GCN attention and binding energy. From this subfigure, we observe the following: (1) The correlation coefficient r = 0.94 indicates that the attention scores learned by the GNN highly align with actual binding energy contributions; (2) Despite its small molecular weight, the hydroxyl group receives high attention (0.88) due to the critical role of hydrogen bonding; (3) The thiazole ring achieves the highest attention (0.92) attributable to its large hydrophobic surface area and shape complementarity. This energy-attention comparison validates the physical interpretability of our model.

Overall, Fig 6 demonstrates that GSAP-GCN can automatically identify key pharmacophores through attention mechanisms without requiring manual feature engineering. Furthermore, the finding is consistent with pharmaceutical facts of the interaction between Ritonavir and HIV-1 protease in the medical field, as shown in Fig 7. It is precisely one of the core advantages of T-pGNN4DTI for DTI prediction.

Influence of the pooling rate in T-pGNN4DTI

The pooling ratio is one of the most important parameters of the global self-attentive pooling layer, which determines the number of nodes retained by the compound molecules in the DTI prediction task and can directly affect the final performance of the T-pGNN4DTI model. To identify the optimal pooling ratio, we evaluated the model performance influence of ten distinct pooling ratios (ranging from 0.1 to 1.0) on the most widely used benchmark dataset, BindingDB.

Fig 8 illustrates the AUC and PRC values predicted by T-pGNN4DTI when predicting drug-target interactions using different pooling ratios on the BindingDB dataset. It can be observed that both AUC and PRC values show an increasing trend as the pooling ratio increases, with the model achieving optimal performance at a pooling ratio of 0.8. This experimental result reveals that all atoms within the drug molecule have a specific contribution to the drug-protein target interaction. Although assigning weights to the nodes can distinguish the contributions of different atoms and thus improve the performance of the prediction model, those atoms with less attention cannot be completely ignored.

thumbnail
Fig 8. The performance influence of different pooling rates. Different pooling rates resulted in significant differences in both AUC and PRC values of the corresponding model.

https://doi.org/10.1371/journal.pone.0352250.g008

Results of the ablation experiment

In this section, the ablation experiments were designed and conducted on the BindingDB and Human datasets to explore the contribution of the global self-attention pooling technique and the PTR pretrained model to the final prediction results of T-pGNN4DTI. First, the global self-attention pooling technique in T-pGNN4DTI is replaced with the most commonly used global max pooling technique in graph convolutional neural networks and BiLSTM, respectively, while the rest of the model structure remains unchanged. This means replacing the proposed model’s global self-attention pooling graph neural network with a standard graph convolutional neural network and BiLSTM to learn drug features, respectively. The corresponding variant models are named T-GCN and T-BiLSTM, respectively. Second, the PTR model in T-pGNN4DTI is replaced with other protein learning models used for learning protein features, including the one-dimensional convolutional neural network (1-D CNN) [40], ESM-1v [51], and ProteinBERT [52], respectively, while the rest of the model structure remains unchanged. These corresponding variant models are named CNN-pGNN, ESM1v-pGNN, and ProBERT-pGNN, respectively. Table 6 demonstrates the comparative results of the ablation experiments. On the benchmark dataset, T-pGNN4DTI demonstrated significantly superior predictive performance across all four evaluation metrics compared to its variants without global self-attention pooling layers and without protein pre-trained models.

CNN-pGNN replaces the PTR model in T-pGNN4DTI with a one-dimensional convolutional neural network, employing GNNs to learn drug features while utilizing CNNs to learn target features. In contrast, T-pGNN4DTI, T-GCN, and T-BiLSTM employ the PTR model to learn protein features and utilize GNNs to learn drug features. CNN-pGNN underperforms the models T-pGNN4DTI and T-GCN on all datasets across all metrics. T-GCN, which replaces the global self-attention pooling graph neural network in T-pGNN4DTI with a conventional GCN to learn drug features, demonstrates significantly poorer performance than T-pGNN4DTI on all datasets across all metrics. These demonstrate that the global self-attention pooling technique and the PTR pretrained model both have a significant contribution to the final prediction results of T-pGNN4DTI.

As shown in Table 6, all comparative models underperformed T-pGNN4DTI across all datasets and metrics. ESM1v-pGNN, ProBERT-pGNN, and T-BiLSTM all underperformed T-pGNN4DTI across all datasets and metrics. This could be attributed to the excellent synergistic complementarity between the PTR and the global attention pooling graph network with the feed-forward neural network in T-pGNN4DTI, enabling its excellent performance in DTI prediction.

Discussion

Our experimental results across three benchmark datasets clearly demonstrate that T-pGNN4DTI achieves overall optimal performance, surpassing existing methods in 10 out of 12 evaluations and establishing a new state-of-the-art. This demonstrates that our model’s architecture strategy—based on a global self-attention pooling graph neural network and pre-trained models—is highly effective.

In comparative experiments, T-pGNN4DTI and T-hGNN4DTI demonstrated superior performance across three benchmark datasets, further validating that DTI prediction methods combining GSAP-GCN with the protein PTR model outperform existing approaches. This advantage stems from the ability of GSAP-GCN and protein PTR models to effectively learn local and global feature information of drugs and proteins, respectively. Furthermore, their synergistic complementarity efficiently integrates the feature information of drugs and targets within drug-target pairs, comprehensively capturing key biological features in drug-target interactions.

T-pGNN4DTI outperformed or equaled T-hGNN4DTI in 11 out of 12 evaluations, indicating that the GCN incorporating global self-attention pooling layers in T-pGNN4DTI learns drug features more effectively than the GCN with hierarchical self-attention pooling layers. Furthermore, T-pGNN4DTI and T-hGNN4DTI outperform T-GCN across all 12 evaluations, as shown in Table 6, suggesting that GCNs incorporating self-attention pooling layers are more effective than those without such layers in molecular graph feature learning scenarios.

The result in the ablation experiment, as shown in Table 6, suggests that our model and its variants (replacing the PTR model with ESM-1v[51], and ProteinBERT [52]) leverage pre-trained protein language models to learn protein representations and their features. This approach outperforms methods such as 1-D CNN and BiLSTM, thereby improving DTI prediction accuracy. This finding strongly aligns with the results from FusionDTI [48] and DTI-LM [33], both of which employ protein pre-trained models to learn protein features for DTI prediction.

Overall, T-pGNN4DTI integrates a global self-attention graph-based neural network and a pre-trained model for DTI prediction, which provides a new way of thinking for solving other DTI-related problems.

However, the benchmarking datasets are not ideal due to limitations in data collection methods and experimental conditions. For example, most drugs in the Human and BindingDB dataset only occur in one class, potentially introducing bias into feature learning. The relatively small sample size of compound-protein interactions (CPIs) in Human and C.elegans may limit the training of complex deep learning models. Furthermore, negative samples are generated by algorithms that may introduce undetectable noise [35], necessitating biological experimental corrections by the scientific community.

Conclusion

In this work, we proposed a DTI prediction method based on a self-attention pooled graph neural network and a pre-trained model, named T-pGNN4DTI. T-pGNN4DTI uses a global self-attention graph convolutional neural network to learn the drug features and uses a pre-trained model based on Transformer to learn the protein features. Experimental comparison results based on three benchmark datasets show that the proposed T-pGNN4DTI model outperforms state-of-the-art methods. The experimental comparison results also demonstrate that using global self-attention pooling graph neural networks and protein pre-trained models to effectively learn drug and target feature information, respectively, is an effective method for improving DTI prediction.

Although the proposed model achieves high predictive performance, its lack of interpretability impedes the understanding of underlying drug–target interaction mechanisms. In addition, it does not fully explore physical constraint information about molecules, which may lead to expertise bias in practical applications and affect its generalization capability. In future research, we will explore methods to enhance model interpretability and incorporate physical constraints into dynamic thermal analysis prediction models to further improve their predictive capabilities.

Supporting information

S1 File. The detailed training strategies of the PTR model.

https://doi.org/10.1371/journal.pone.0352250.s001

(DOCX)

S2 File. Time Complexity and Computational Efficiency.

https://doi.org/10.1371/journal.pone.0352250.s002

(DOCX)

References

  1. 1. Landry Y, Gies J-P. Drugs and their molecular targets: an updated overview. Fundam Clin Pharmacol. 2008;22(1):1–18. pmid:18251718
  2. 2. Shi W, Yang H, Xie L, Yin X-X, Zhang Y. A review of machine learning-based methods for predicting drug-target interactions. Health Inf Sci Syst. 2024;12(1):30. pmid:38617016
  3. 3. Abdul Raheem AK, Dhannoon BN. Comprehensive Review on Drug-target Interaction Prediction - Latest Developments and Overview. Curr Drug Discov Technol. 2024;21(2):e010923220652. pmid:37680152
  4. 4. Liao Q, Zhang Y, Chu Y, Ding Y, Liu Z, Zhao X. Application of Artificial Intelligence in Drug-target Interactions Prediction: A Review. npj Biomedical Innovations. 2025;2(1):1–19.
  5. 5. Zhang R, Lin Y, Wu Y, Deng L, Zhang H, Liao M, et al. MvMRL: a multi-view molecular representation learning method for molecular property prediction. Brief Bioinform. 2024;25(4):bbae298. pmid:38920342
  6. 6. Peng Y, Zhang Z, Jiang Q, Guan J, Zhou S. TOP: A deep mixture representation learning method for boosting molecular toxicity prediction. Methods. 2020;179:55–64. pmid:32446957
  7. 7. Tanoli Z, Schulman A, Aittokallio T. Validation guidelines for drug-target prediction methods. Expert Opin Drug Discov. 2025;20(1):31–45. pmid:39568436
  8. 8. Wang B, Zhang T, Liu Q, Sutcharitchan C, Zhou Z, Zhang D, et al. Elucidating the role of artificial intelligence in drug development from the perspective of drug-target interactions. J Pharm Anal. 2025;15(3):101144. pmid:40099205
  9. 9. Li D, Zhao F, Yang Y, Cui Z, Hu P, Hu L. Multi-View Contrastive Learning for Drug-Drug Interaction Event Prediction. IEEE J Biomed Health Inform. 2026;30(3):2276–87. pmid:40833904
  10. 10. Zheng W, Li Z, Chen Y, Liao W, Deng L, Zhang H, et al. GEP-DNN4Mol: automatic chemical molecular design based on deep neural networks and gene expression programming. Health Inf Sci Syst. 2025;13(1):31. pmid:40144479
  11. 11. Abbasi K, Razzaghi P, Poso A, Ghanbari-Ara S, Masoudi-Nejad A. Deep Learning in Drug Target Interaction Prediction: Current and Future Perspectives. Curr Med Chem. 2021;28(11):2100–13. pmid:32895036
  12. 12. Zheng X, Ding H, Mamitsuka H, Zhu S. Collaborative matrix factorization with multiple similarities for predicting drug-target interactions. In: Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining, 2013. 1025–33.
  13. 13. Cichonska A, Ravikumar B, Parri E, Timonen S, Pahikkala T, Airola A, et al. Computational-experimental approach to drug-target interaction mapping: A case study on kinase inhibitors. PLoS Comput Biol. 2017;13(8):e1005678. pmid:28787438
  14. 14. He T, Heidemeyer M, Ban F, Cherkasov A, Ester M. SimBoost: a read-across approach for predicting drug-target binding affinities using gradient boosting machines. J Cheminform. 2017;9(1):24. pmid:29086119
  15. 15. Wen M, Zhang Z, Niu S, Sha H, Yang R, Yun Y, et al. Deep-Learning-Based Drug-Target Interaction Prediction. J Proteome Res. 2017;16(4):1401–9. pmid:28264154
  16. 16. Peng J, Li J, Shang X. A learning-based method for drug-target interaction prediction based on feature representation learning and deep neural network. BMC Bioinformatics. 2020;21(Suppl 13):394. pmid:32938374
  17. 17. Nguyen T, Le H, Quinn TP, Nguyen T, Le TD, Venkatesh S. GraphDTA: predicting drug-target binding affinity with graph neural networks. Bioinformatics. 2021;37(8):1140–7. pmid:33119053
  18. 18. Zhao B-W, Su X-R, Hu P-W, Huang Y-A, You Z-H, Hu L. iGRLDTI: an improved graph representation learning method for predicting drug-target interactions over heterogeneous biological information network. Bioinformatics. 2023;39(8):btad451. pmid:37505483
  19. 19. Zhang Y, Hu Y, Han N, Yang A, Liu X, Cai H. A survey of drug-target interaction and affinity prediction methods via graph neural networks. Comput Biol Med. 2023;163:107136. pmid:37329615
  20. 20. Veličković P, Cucurull G, Casanova A, Romero A, Liò P, Bengio Y. Graph Attention Networks. In: 2018. https://openreview.net/forum?id=rJXMpikCZ
  21. 21. Yazdani-Jahromi M, Yousefi N, Tayebi A, Kolanthai E, Neal CJ, Seal S, et al. AttentionSiteDTI: an interpretable graph-based model for drug-target interaction prediction using NLP sentence-level relation classification. Brief Bioinform. 2022;23(4):bbac272. pmid:35817396
  22. 22. Song W, Xu L, Han C, Tian Z, Zou Q. Drug-target interaction predictions with multi-view similarity network fusion strategy and deep interactive attention mechanism. Bioinformatics. 2024;40(6):btae346. pmid:38837345
  23. 23. Talo M, Bozdag S. Top-DTI: integrating topological deep learning and large language models for drug-target interaction prediction. Bioinformatics. 2025;41(Supplement_1):i133–41. pmid:40662785
  24. 24. Abbasi K, Razzaghi P, Poso A, Amanlou M, Ghasemi JB, Masoudi-Nejad A. DeepCDA: deep cross-domain compound-protein affinity prediction through LSTM and convolutional neural networks. Bioinformatics. 2020;36(17):4633–42. pmid:32462178
  25. 25. Bai P, Miljković F, John B, Lu H. Interpretable bilinear attention network with domain adaptation improves drug–target prediction. Nature Machine Intelligence. 2023;5(2):126–36.
  26. 26. Zhao Q, Xiao F, Yang M, Li Y, Wang J. AttentionDTA: prediction of drug–target binding affinity using attention model. In: 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2019. 64–9. https://doi.org/10.1109/bibm47256.2019.8983125
  27. 27. Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: 2019. https://doi.org/10.48550/arXiv.1810.04805
  28. 28. Li D, Yang Y, Cui Z, Yin H, Hu P, Hu L. LLM-DDI: Leveraging Large Language Models for Drug-Drug Interaction Prediction on Biomedical Knowledge Graph. IEEE J Biomed Health Inform. 2026;30(1):773–81. pmid:40601466
  29. 29. Kalakoti Y, Yadav S, Sundar D. TransDTI: Transformer-Based Language Models for Estimating DTIs and Building a Drug Recommendation Workflow. ACS Omega. 2022;7(3):2706–17. pmid:35097268
  30. 30. Aldahdooh J, Vähä-Koskela M, Tang J, Tanoli Z. Using BERT to identify drug-target interactions from whole PubMed. BMC Bioinformatics. 2022;23(1):245. pmid:35729494
  31. 31. Kang H, Goo S, Lee H, Chae J-W, Yun H-Y, Jung S. Fine-tuning of BERT Model to Accurately Predict Drug-Target Interactions. Pharmaceutics. 2022;14(8):1710. pmid:36015336
  32. 32. Tang W, Zhao Q, Wang J. LLMDTA: Improving Cold-Start Prediction in Drug-Target Affinity With Biological LLM. IEEE Trans Comput Biol Bioinform. 2025;22(6):2398–409. pmid:40811267
  33. 33. Ahmed KT, Ansari MI, Zhang W. DTI-LM: language model powered drug-target interaction prediction. Bioinformatics. 2024;40(9):btae533. pmid:39221997
  34. 34. Chen L, Tan X, Wang D, Zhong F, Liu X, Yang T, et al. TransformerCPI: improving compound-protein interaction prediction by sequence-based deep learning with self-attention mechanism and label reversal experiments. Bioinformatics. 2020;36(16):4406–14. pmid:32428219
  35. 35. Liu H, Sun J, Guan J, Zheng J, Zhou S. Improving compound-protein interaction prediction by building up highly credible negative samples. Bioinformatics. 2015;31(12):i221-9. pmid:26072486
  36. 36. Liu T, Lin Y, Wen X, Jorissen RN, Gilson MK. BindingDB: a web-accessible database of experimentally determined protein-ligand binding affinities. Nucleic Acids Res. 2007;35(Database issue):D198-201. pmid:17145705
  37. 37. Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci U S A. 2021;118(15):e2016239118. pmid:33876751
  38. 38. Mikolov T, Sutskever I, Chen K, Corrado G, Dean J. Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems. 2013;26.
  39. 39. Lee J, Lee I, Kang J. Self-Attention Graph Pooling. In: Proceedings of the 36th International Conference on Machine Learning, 2019. https://doi.org/arXiv.1904.08082
  40. 40. Zhang S, Jiang M, Wang S, Wang X, Wei Z, Li Z. SAG-DTA: Prediction of Drug-Target Affinity Using Self-Attention Graph Network. Int J Mol Sci. 2021;22(16):8993. pmid:34445696
  41. 41. Zhao M, Yuan M, Yang Y, Xu SX. CPGL: Prediction of Compound-Protein Interaction by Integrating Graph Attention Network With Long Short-Term Memory Neural Network. IEEE/ACM Trans Comput Biol Bioinform. 2023;20(3):1935–42. pmid:36445995
  42. 42. Wang L, Zhou Y, Chen Q. AMMVF-DTI: A Novel Model Predicting Drug-Target Interactions Based on Attention Mechanism and Multi-View Fusion. Int J Mol Sci. 2023;24(18):14142. pmid:37762445
  43. 43. Qian Y, Li X, Wu J, Zhang Q. MCL-DTI: using drug multimodal information and bi-directional cross-attention learning method for predicting drug-target interaction. BMC Bioinformatics. 2023;24(1):323. pmid:37633938
  44. 44. Li W, Zhao J, Zhu L, Zhang L, Wang H. TopoPharmDTI: Improving interactions prediction by enhanced deep learning representation for both drug and target molecules. Tsinghua Science and Technology. 2026;31(1):399–417.
  45. 45. Zhang J, Liu Z, Pan Y, Lin H, Zhang Y. IMAEN: An interpretable molecular augmentation model for drug–target interaction prediction. Expert Systems with Applications. 2024;238(PC):10.
  46. 46. Peng L, Liu X, Chen M, Liao W, Mao J, Zhou L. MGNDTI: A Drug-Target Interaction Prediction Framework Based on Multimodal Representation Learning and the Gating Mechanism. J Chem Inf Model. 2024;64(16):6684–98. pmid:39137398
  47. 47. Qian Y, Wang Q, Yin L, Lu A-Y. CF-DTI: coarse-to-fine feature extraction for enhanced drug-target interaction prediction. Health Inf Sci Syst. 2025;13(1):55. pmid:40919585
  48. 48. Meng Z, Meng Z, Yuan K, Ounis I. FusionDTI: Fine-grained Binding Discovery with Token-level Fusion for Drug-Target Interaction. In: Findings of the Association for Computational Linguistics: EMNLP 2025, 2025. 4425–44. https://doi.org/10.18653/v1/2025.findings-emnlp.237
  49. 49. Tsubaki M, Tomii K, Sese J. Compound-protein interaction prediction with end-to-end learning of neural networks for graphs and sequences. Bioinformatics. 2019;35(2):309–18. pmid:29982330
  50. 50. Zheng S, Li Y, Chen S, Xu J, Yang Y. Predicting drug–protein interaction using quasi-visual question answering system. Nature Machine Intelligence. 2020;2(2):134–40.
  51. 51. Meier J, Rao R, Verkuil R, Liu J, Sercu T, Rives A. Language models enable zero-shot prediction of the effects of mutations on protein function. In: NIPS ’21, 2021.
  52. 52. Brandes N, Ofer D, Peleg Y, Rappoport N, Linial M. ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics. 2022;38(8):2102–10. pmid:35020807