Figures
Abstract
Generative machine learning models offer a powerful framework for therapeutic design, by efficiently exploring large spaces of biological sequences enriched for desirable properties. Unlike supervised learning methods, which require both positive and negative labeled data, generative models such as LSTMs can be trained solely on positively labeled sequences, for example, high-affinity antibodies. This is particularly advantageous in biological settings where negative data are scarce, unreliable, or biologically ill-defined. However, the lack of attribution methods for generative models has hindered the ability to extract interpretable biological insights from such models. To address this gap, we developed Generative Attribution Metric Analysis (GAMA), an attribution method for autoregressive generative models based on Integrated Gradients. We assessed GAMA using synthetic datasets with known ground truths to characterize its statistical behavior and validate its ability to recover biologically relevant features. We further demonstrated the utility of GAMA by applying it to experimental antibody-antigen binding data. GAMA enables model interpretability and the validation of generative sequence design strategies without the need for negative training data.
Author summary
We have developed a new way to look inside AI language models that create biological sequences such as antibodies, even when only examples of successful binders are available. Traditional design methods need both positive and negative examples, but in many biomedical settings negative data are scarce, unreliable, or undefined. By training a model on only the desirable (positive) sequences and then comparing the model’s internal “gradient” signals to those of an untrained reference, we can identify which positions in a protein sequence the model deems important for binding. This approach, which we call Generative Attribution Metric Analysis (GAMA), lets us recover known binding motifs from synthetic data, confirm biologically relevant positions in simulated antibody‑antigen interactions, and highlight key residues in real experimental antibody datasets. The method therefore adds a layer of interpretability to otherwise “black-box” AI-based protein design tools, helping researchers understand why a generated sequence is predicted to work and enabling more confident use of these models in therapeutic development and other areas where only positive examples exist.
Citation: Frank R, Widrich M, Akbar R, Klambauer G, Sandve GK, Robert PA, et al. (2026) Attribution assignment for deep-generative sequence models enables interpretability analysis using positive-only data. PLoS Comput Biol 22(9): e1014805. https://doi.org/10.1371/journal.pcbi.1014805
Editor: Pingzhao Hu, Western University, CANADA
Received: August 23, 2025; Accepted: September 7, 2026; Published: September 28, 2026
Copyright: © 2026 Frank et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All code is available at https://github.com/csi-greifflab/GAMA.
Funding: This was supported by a UiO: LifeScience Convergence Environment Immunolingo (to VG and GKS), a Norwegian Cancer Society Grant (#215817, to VG), Research Council of Norway projects (#300740, #331890 to VG), a Research Council of Norway IKTPLUSS project (#311341, to VG and GKS), Stiftelsen Kristian Gerhard Jebsen (K.G. Jebsen Coeliac Disease Research Centre) (to GKS) and the European Union (ERC, AB-AG-INTERACT, 101125630, to VG). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: I have read the journal’s policy and the authors of this manuscript have the following competing interests: V.G. declares advisory board positions in aiNET GmbH, Enpicom B.V, Absci, Omniscope, and Diagonal Therapeutics. V.G. is a consultant for Adaptive Biosystems, Specifica Inc, Roche/Genentech, immunai, LabGenius, and FairJourney Biologics. V.G. is an employee of Imprint LLC. PR declares employment by F. Hoffmann-La Roche AG.
Introduction
Deep generative neural networks are a class of models capable of learning complex data distributions and can enable the generation of novel instances from them. They do not require labels and can be trained on “positive-only samples” (hereafter referred to “1-class samples”). Some of the most common approaches used in generative modeling are Long Short-Term Memory networks (LSTM) [1], state space models [2], Generative Adversarial Networks (GAN) [3], Variational AutoEncoder (VAE) [4], diffusion models [5], and transformer networks [6].
Several protein design approaches have succeeded using (deep) generative models [7–19]. Generative models have also been used to design DNA and RNA sequences, e.g., to design novel promoter sequences or RNA sequences with desired folding properties [20–23]. Previous reports have demonstrated that monoclonal antibodies can be synthetically generated using generative models and have a distribution of binding affinity and various biophysical parameters that is similar to the training data distribution [7,8,24–30].
The interpretability of deep learning methods is inherently challenging, as they process information in highly nonlinear hierarchical representations. Various attribution assignment methods have been developed to address this issue, aiming to elucidate the contribution of input features to the output decision [31–33]. These methods are instrumental in diagnosing individual classification errors and assessing the confidence level of model reliability [34]. Among these techniques, Integrated Gradients (IG) stands out as a particularly effective approach [35]. IG emphasizes the significance of input features in the prediction process and has been successfully employed in various supervised sequence prediction tasks. For instance, Senior and colleagues demonstrated IG’s utility in elucidating the influence of feature location on protein folding prediction from sequences [36,37]. Similarly, Preuer and colleagues showcased IG’s role in molecule design for drug discovery [38], and Ishida and colleagues utilized IG to identify key atoms in molecules for reaction prediction [39].
The use of IG in supervised LSTM models is well-documented across multiple studies [40,41], including its application to DNA sequences [42] and RNA sequences [43]. IG has also been implemented in the analysis of antibody sequences: our recent work revealed that IG, when employed in binary classification methods, uncovers structural patterns relevant to antibody binding [44,45]. Furthermore, Widrich and colleagues identified pertinent patterns in amino acid sequences within immune repertoire datasets using IG [46]. A detailed guide on integrating the attribution assignment method “Layer-wise Relevance Propagation” into LSTM models was provided by Arras and colleagues [47].
Extensive research has been performed on rendering supervised methods interpretable. However, there is only limited research available that has focused on the interpretability of generative sequence models [48,49], especially generative sequence models for 1-class data [50]. There are several papers on developing attribution assignment methods for transformers and large language models [51–54]. There exists some work on interpreting generative (sequence) models with a latent space [55–57]. Approaches that tackle attribution assignment methods for 1-class generative models are rare [57,58]. A survey paper from Tritscher et al. is available [59]. To our knowledge, there exists no prior work which uses gradient-based methods on LSTMs for attempting to identify essential patterns that are relevant for LSTM-based sequence generation.
Interpretable methods would enable addressing the negative label bias problem, which primarily occurs in immunological datasets but poses a challenge in many supervised multiclass settings where the definition of negative data is ambiguous [44,45,60–62]. Specifically, supervised methods for immune receptor classification tasks require at least two labels in the training data, where one class usually consists of those receptor sequences that have been validated to bind the target antigen. The other, negative class, is challenging to define as it may be defined as non-binders or, for example, binders to another target antigen [45]. Indeed, we and others have previously shown that the training data composition strongly impacts the classification of the positive class by ML methods and best practices for defining negative datasets are still controversial [44,45,60–62]. Applying attribution assignment methods to solely 1-class (positive‑only) data (i.e., antigen-specific binders) using generative models would reduce the requirements for training data size and complexity, enabling the generation of both novel binders and an understanding of the underlying binding patterns. Furthermore, interpretable generative models would enable greater trust and confidence in sequences designed by such models, making them more viable for sensitive applications, such as drug design [63,64].
To elucidate the underlying determinants and biological principles that influence sequence generation by autoregressive generative models (Fig 1A), we have developed a novel attribution assignment method (GAMA) (Fig 1B), which highlights the importance of specific positions for the positive class. GAMA is based on Integrated Gradients (IG) for LSTMs. We demonstrated GAMA’s application to LSTMs trained on synthetic sequences and the results suggest that the importance of features gives us information on binding rules, i.e., amino acids that are important for binding. We define “Importance” as the difference between the integrated gradients (attribution assignment method) of a trained LSTM and a randomly initialized LSTM. We also show the applicability to antibody sequences.
(A) Deep generative sequence models enable the design of novel biological sequences from only positively labeled datasets, such as antibody CDR sequences. However, deep generative models are black box models and it is unclear what patterns they learn. We investigate how to interrogate the learned distribution of generative sequence models and how generative sequence models can be used to understand the training distribution. (B) We developed a method called “Generative Attribution Metric Analysis” (GAMA) to identify learned motifs in generative models and make them more interpretable. Generative models only require 1-class data that consist solely of positive instances. For example, antibody binding datasets have only instances of antigen-binding (positive) sequences available but it is difficult to experimentally acquire sequences that explicitly do not bind the antigen. (C) We evaluate GAMA on 270 synthetic sequence datasets, four synthetic antibody sequence datasets, and one experimental antibody sequence dataset to demonstrate retrieval efficiency and interpretability.
Briefly, LSTMs and other state-space models estimate a probability distribution for the next token in a sequence given the previous tokens [65]. In this way, LSTMs model the entire sequence distribution globally and unsupervised while learning a supervised model locally for next-token prediction [66]. IG works on the local level of the autoregressive model, i.e., predicting a single element from the previous element, by identifying the impact of one input-output token pair. By summarizing the IG results over all possible combinations of input-output pairs, GAMA obtains the global distribution that the model has learned.
Monoclonal antibodies (mAb) are used for various applications, such as therapeutics or as diagnostic biomarkers [63]. Each individual drug and application requires a specifically designed mAb. MAb design is currently performed both in vivo or in vitro [67]. This is usually very time and cost-intensive, for example, lead times to mAb discovery and design are on average >3 years [64, 68]. Generative in-silico design of mAb sequences is a promising avenue for improving mAb design lead times and costs. Additionally, mAb sequences are a good modeling [44] system for evaluating our methods due to the availability of synthetic ground truth data [44]. The structural binding of an antibody to a target protein is described by a set of amino acids (termed paratopes) that interact with the binding site of the target antigen (termed epitopes) [63]. Paratopes are typically localized in the complementarity-determining regions (CDR) of antibodies. Paratopes can have complex interactions with epitopes [63]. Consequently, attribution assignment methods must strongly consider the complexity [44] of the datasets and the robustness towards noise. Paratopes are usually gapped motifs [69] and we investigate the hypothesis that attribution methods will focus on binding residues preferentially, an analysis of attributions could be used to predict the paratope of antibody sequences. Successful attribution assignment of these motifs will pave the way for interpretable deep generative models [70].
In this work, we trained an LSTM with an autoregressive unsupervised objective on synthetic 1-class CDR3-antigen binding data at different levels of complexity as well as real-world data and identified the important features learned by the generative model using our newly developed method “Generative Attribution Metric Analysis” (GAMA) (Fig 1C).
Results
A method to quantify and highlight important motifs in generative models
Generative models have been used to learn and generate antibody sequences with desired properties. We develop a tool, GAMA, to extract information from trained generative models (Fig 2A). In particular, we identify the features learned by a generative model as follows: First, we apply interpretability/attribution approaches, such as IG, to the trained autoregressive generative model “main model” to obtain the importance of each input/output pair - e.g., the LSTM predicting the next amino acid based on the previous ones (Fig 2B). Subsequently, we repeat this process for a “reference model” and use the difference between the main and the reference model to obtain the learned features that are particular to the main model. See Algorithm 1 for details.
(A) GAMA highlights important sequence positions that were learned by a generative LSTM model. We define “Importance” as the difference between the integrated gradients (attribution assignment method) of a trained LSTM and a randomly initialized LSTM. (B) The gradients can be computed for every output token at each sequence position w.r.t the input encoding. The figure illustrates the computation of gradients for the fourth “C” output; the red arrows indicate the computation flow of the gradients and the red tensors represent the gradient components of each input position. We compute the gradient components for several encodings that are obtained by linearly interpolating between the encoding of the input sequence and a “zero input baseline”. The gradient components are pooled together to compute a global integrated gradient. They are computed for a trained LSTM with a sequence generated from the same LSTM over all outputs. GAMA consolidates all possible gradients into a statistic that highlights important sequence positions. (C) GAMA is evaluated on three distinct datasets with increasing levels of real-world applicability for antibody binding (270 synthetic datasets, Absolut!-simulated datasets, and Trastuzumab-HER2 binder dataset). We test our method by applying GAMA to models trained on datasets with known features.
In the context where an antibody dataset contains binding sequences to an antigen, we hypothesize that if a generative model has learnt biological rules of antibodies, GAMA should reveal AAs important for the binding (Fig 2), and define the capacity to identify motifs as the predictive power of the model. We define a motif as a set of repeating AAs at the same positions. We test our method by applying GAMA to models trained on datasets with known binding patterns. Since datasets with known binding patterns are rare, we first use 270 simple synthetic datasets with specific properties and implanted motifs to assess motif recovery under noise, motif position, inter-motif interactions, and motif length. We define motifs as short repeating elements of a sequence in a dataset. We assume this simulates important features of biological datasets that all share a single common gapped motif (paratope motifs are usually gapped motifs [69, 71]). Then we evaluate GAMA on complex simulation-derived Ab-Ag binding simulations [44] with known binding motifs to assess if generative models can learn binding motifs. Lastly, we evaluate GAMA on an experimental dataset of HER-2 binding antibody sequences [72] (the Trastuzumab paratope is a non-gapped motif of size four) where we are only certain of the motif of one of the sequences (Fig 2C).
An LSTM is an autoregressive generative sequence model, which means it iteratively predicts the next element within the sequence given all previous sequences (next-character prediction), and is trained with teacher forcing. Such a model can sample novel sequences given a “start token”.
GAMA is derived by comparing IG signals from a randomly initialized reference LSTM and a trained LSTM model; the reference model uses exactly the same weights that were later used for training, so that the only systematic difference between the two models is the effect of exposure to the true data distribution.
The importance metric (GAMA) is derived by comparing IG signals from a randomly initialized reference LSTM and a trained LSTM model. The reference model is identical to the initialization used to train the model thereby giving us the difference the training data has induced into the model. Our similarity function M(t: = token position) is described in Equation 1. The similarity function measures the absolute median difference between the IG values for the reference model and the trained model (Fig 3). The difference was weighted by the sum of the variance of the IG values in the reference model and the trained model and normalized by the total variance of the sequence.The metric was inspired by the two empirical observations of signal‑position behavior.
Relevant signal motifs can be identified by comparing the IG of a trained model and the IG of a randomly initialized model. Here we use an LSTM to showcase the capability of our metric to assign attribution in autoregressive generative models. (1) An LSTM is initialized randomly and saved as the “reference model”. (2) The reference model is trained generatively on antibody (Ab) sequences. The resulting model is then saved as “trained model”. (3) For both models, we computed the IG-tensor (Fig 10) and reduced the dimension from 4D to 2D tensor. (4) We computed the similarity function for the two tensors (Equation 1).
- They exhibit a large absolute median shift between the trained and reference IG vectors.
- IG values at those positions are more dispersed (higher variance) than at background positions.
By explicitly modelling both aspects, Equation 1 captures the complete signature of a signal. The numerator is an estimate of the effect, while the denominator estimates the relative noise. Although the metric was initially motivated by the observed median‑difference/variance pattern; it follows a well‑established statistical recipe for comparing two empirical distributions: robust central‑tendency difference divided by a pooled dispersion measure.
In the next section, we will benchmark GAMA on a simple synthetic dataset, a more challenging antibody synthetic data [44], and antibody experimental data [72].
Equation 1 describes the computation of the similarity function M(s) that computes the similarity of our trained model IG distribution and our reference model IG distribution at token position t. indicates all IG datapoints computed on the trained model at sequence token position t. N indicates the total sequence length. A large difference indicates an important sequence position.
Benchmarking of LSTM-based motif recovery across three key parameters: noise ratio, motif position, and motif complexity on synthetic datasets
We investigated how GAMA is affected by three variables: (i) noise ratio (ii) motif position (iii) motif logic. Noise ratio (i), is a ratio defined by the number of sequences containing motifs (motif_sequ) and the number of contaminating sequences not containing any motifs (no_motif):
Motif position (ii) refers to whether the amino acids of a motif are concentrated towards the beginning, middle, or end of a sequence. Motif logic (iii) designates how the amino acids within a motif are allowed to interact to be counted as a motif. Specifically, motif interaction is a function of all motif positions; each motif position either has the relevant AA (true) or not (false) if the function returns true then the sequence is considered to be part of the dataset. For example, a motif with “AND” logic must have all associated positions be set to true to be counted as a motif. We investigate three logic types: “AND”, “OR”, and exclusive or “XOR”.
To investigate the impact of these three parameters (parameter 3-tuples), we explored all possible combinations of the parameter 3-tuples. Each experiment consisted of a synthetic dataset as described in Methods. Briefly, each sequence had a length of 16, a value chosen because it approximately represents the median length of human CDRH3 sequences [73].the 3-tuples had the following values: three different logic motif interactions, three motif positions times three motif lengths, and ten noise conditions, which resulted in 3*(3*3)*10 = 270 experiments (Fig 2C). An example of such a 3-tuple is a “L” implanted at position two, a “Y” implanted at position four, with a motif logic of “OR”, and a signal ratio of 80% as seen in Fig 5. Since we implanted the artificial motif into these datasets, we know the ground truth motif position.
We investigated the usefulness of the GAMA importance values by applying it to LSTM models that were trained on datasets with known features and utilizing the importance values in a motif retrieval task. We demonstrated sufficient model convergence by sampling from the resulting models and plotting their frequency distribution (S21 Fig). We quantify GAMA performance by measuring the overlap between the top‑ranked positions (according to GAMA) and the known ground‑truth motif positions. Retrieval performance is the ability of GAMA to recover relevant motif sequences accurately. Therefore, high retrieval performance indicates high interpretability. The noise ratio tests whether GAMA is robust to noise within the dataset; in biological datasets noise can be introduced, for example, through imperfect FACS sorting or library amplification or sequencing-related errors. The motif position parameter investigates if the retrieval performance is position invariant because it could be degraded due to gradient vanishing. The motif logic investigates the ability of our method to identify motifs with non-linear amino acid interactions.
For comparison, we define an “IG baseline” that summarizes Integrated Gradients computed only on the trained LSTM, without reference to a randomly initialized model. Concretely, for each sequence position t, we aggregate the absolute IG values over all input token types, all output token types, and all output positions, yielding a single baseline importance score per position. Unlike GAMA, which explicitly contrasts trained and reference IG distributions to highlight training-induced changes, the IG baseline reflects a conventional, reference-free attribution derived solely from the trained model.
GAMA achieves high motif retrieval performance at high signal ratios
We sought to understand how robust GAMA is in noisy datasets. Biological dataset creation is error-prone, i.e., sequences that are wrongly labeled as positive [74]. For example, FACS cell sorting can introduce noise when identifying high and low binding Ab binders, thus, introducing false positives into the dataset, we are not considering false negatives because we do not use negatively labeled data. Large numbers of sequences without motifs can also be inherent to certain datasets, for example, antibody repertoires sample all immune sequences from a subject with a disease status, but each subject will have many non-disease status-associated antigens.
To assess how the presence of spurious (false‑positive) sequences influences motif recovery, we systematically introduced uniformly random amino‑acid sequences into the training set at varying proportions, which we refer to as the signal ratio. The signal ratio denotes the fraction of genuine signal sequences relative to the total dataset size. Starting from a fixed pool of 10 000 signal sequences, a signal ratio of 100% corresponds to a dataset composed exclusively of these sequences (no noise), whereas a signal ratio of 10% implies that only 10% of the final dataset (1 000 sequences) are true signal and the remaining 90% (9 000 sequences) are random noise. Fig 5 illustrates an example dataset constructed with a signal ratio of 80% (8 000 signal sequences and 2 000 noise sequences).
We increase the signal ratio from 0% to 100% in increments of 10 percentage points. For each signal ratio we consider three separate motif-logic conditions (AND, OR, XOR) at each 9 different motif positions. This results in a total of 270 experiments (Fig 2C). In Fig 4 we show the results for all experiments with signal positions two and four for all motif logic types without averaging over them as an example. We can see that the GAMA attributions have the strongest distinction on the two and four positions for a signal ratio of 100% (no noise sequences) and “AND” motif logic. This distinction diminishes with more noise sequences and for “OR” and “XOR” logic datasets. We also created the same visualization with an IG baseline for comparison in S3 Fig. Further examples for GAMA and IG baseline for different motifs can be found in S4, S5, S6, S7, S8, S9, S10, and S11 Figs. In S24 Fig we show a synthetic experiment repeated with a different seed. Additionally, we assessed how well this synthetic data can be classified by the classical positive‑only algorithm (Isolation Forest [75]) S17 Fig. Because Isolation Forest operates at the sequence level and only distinguishes between noise and non‑noise sequences, its outputs are not directly comparable to the motif retrieval performance of our GAMA method. The Isolation Forest baseline achieves mostly close to random balanced accuracy, especially for longer motifs, indicating that sequence‑level positive‑only classification is mostly unable to leverage the implanted motifs. For comparison, we evaluated the performance of standard Integrated Gradients applied to LSTM models trained in a supervised manner on the same synthetic dataset S16 Fig.
This figure shows the GAMA importance values for datasets of uniformly random amino acid sequences with implanted motifs at positions 2 and 4. Each dataset contains 10 000 sequences of length sixteen. In total, 30 datasets were generated, covering all combinations of motif logic (AND, OR, XOR) and signal ratios. The figure is a facet plot of bar charts. Each facet corresponds to a specific combination of motif logic (columns) and signal ratio (rows). Within each facet, bars show GAMA importance values (as defined in the Methods) for all sequence positions on a logarithmic scale, with sequence position on the x‑axis and GAMA magnitude on the y‑axis. Across conditions, GAMA places greater importance on the true motif positions 2 and 4, particularly at high signal ratios. As the proportion of noise sequences increases, GAMA values at the motif positions decrease, while values at non‑motif positions remain relatively constant, indicating that GAMA is sensitive to the implanted motif signal and degrades gracefully with increasing noise. S3 Fig depicts the same visualization for the IG baseline.
The results of this supervised baseline are directly comparable to the GAMA retrieval performance; however, the task it solves is trivial, so its superior performance relative to GAMA is unsurprising. Nevertheless, it highlights motif and logic configurations in the synthetic dataset that remain challenging even for supervised approaches, providing a useful reference for interpreting GAMA’s limitations. Supervised LSTM classifiers combined with Integrated Gradients perfectly recover motif positions under AND and, to a slightly lesser extent, OR logic, but still largely fail to correctly attribute positions under XOR logic. This pattern mirrors our GAMA results, where motif retrieval is strong for AND and OR conditions but substantially weaker for XOR, demonstrating the difficulty of capturing XOR‑type interactions for both supervised and unsupervised/one-class attribution methods.
The drop in performance under high noise levels or when motifs follow more complex logical relationships stems from two factors. First, as the proportion of random noise sequences increases, the gradients that Integrated Gradients relies on become increasingly diluted; the signal contributed by genuine motif residues is averaged together with many irrelevant background contributions, lowering the contrast between true‑positive and background positions. Second, complex motif interactions (e.g., OR or XOR logic) create non‑linear dependencies among residues, so the attribution of a single token no longer captures the joint effect of the motif. Consequently, the IG‑derived importance for each individual position is weaker and more variable, leading GAMA to miss or mis‑rank those positions when identifying the underlying pattern.
For every synthetic data set, we know exactly which positions in the sequence constitute the implanted motif; we call this set the ground‑truth motif. After GAMA is applied, each position in the sequence receives an importance score and the positions are ranked from most to least important. We then take the top‑ranked positions equal in number to the size of the ground‑truth motif (for example, if the motif occupies four positions, we examine the four highest‑scoring positions). The false‑negative‑rate is defined as the proportion of true‑motif positions that are missed. It is calculated by dividing the number of false negatives by the total number of motif positions. Consequently, a false-negative rate of 0% means that every true motif position was recovered, whereas a false-negative rate of 90% means that only 10% of the true motif positions were recovered (lower is better).
We averaged the false-negative-rates for 9 position experiments and plotted them with respect to each signal ratio for motifs that were formed by one of the logic conditions (Fig 5B–5D).
(A) We defined 270 experimental conditions, with each experimental condition being a synthetically generated dataset, the sequences are sampled randomly first, each sequence had a length of 16, a value chosen because it approximately represents the median length of human CDRH3 sequences. Then a motif is implanted into the non-noise sequences. Each dataset is generated by varying three key parameters: “Signal Ratio”, “Signal Position”, and “Motif Logic”, as visualized by the cube. An example dataset with OR logic, 80% signal ratio, and (2, 4) signal position is visualized. For each possible combination of parameters, we create one dataset consisting of 10 000 sequences containing the motif and an appropriate amount of noise sequences without the motif, depending on the “Signal Ratio”-condition. Each parameter has an associated false-negative-rate for GAMA and the IG baseline. A general random baseline is computed that indicates the average false-negative-rate if the top-k motifs positions are identified randomly. (B) The false-negative-rate decreases with higher signal to noise ratio for motifs that are connected by AND logic. For signal ratios of ≥80%, we obtained perfect results (0 false-negative-rate). (C) The false-negative-rate decreases with higher signal ratio for motifs that are connected by OR logic. OR-logic motifs perform worse than AND-logic motifs but better than Exclusive-OR-logic. (D) The false-negative-rate is above 50% for all signal ratio conditions ≥80%. The false-negative-rate decreases with higher signal ratio for motifs that are connected by Exclusive-OR-logic (XOR). Motifs connected by XOR-logic are the most challenging to recover and the GAMA results are similar to the IG baseline. The numbers on the bar charts signify the false-negative-rate.
Since, to the best of our knowledge, there are no attribution assignment methods for 1‑class sequence generative models available, we computed a random baseline with a Monte Carlo simulation to derive the false-negative-rate, one would get if the motifs were retrieved at uniform random; resulting in a false-negative-rate 0.93. We also computed the false-negative-rate for the IG baseline. GAMA outperformed the random baseline in all experimental conditions. GAMA had the best performance for motifs with AND-logic, motifs were retrieved with 0% false-negative-rate for a signal ratio ≥80%. For a signal ratio of 40%–70%, the signal is identified with a false-negative-rate lower than 50% (Fig 5B). GAMA could still recover motifs with an OR-logic but the performance was worse for the entire signal ratio range. Motifs were retrieved with a false-negative-rate lower than 50% for signal ratio of 50%–100% (Fig 5C). Motifs with Exclusive-OR logic are the most difficult to retrieve with GAMA. The false-negative-rate is worse for the entire signal ratio range and the false-negative-rate is below 50% only for signal ratios of 90%-100%. To summarize, GAMA was able to identify the signal across all noisy dataset conditions to some extent. We investigated how effectively the GAMA function retrieves the motif positions, for all 270 datasets. Subsequently, we calculated and visualized the false-negative error rate (S12 Fig), and the TopK (until full retrieval) (S13 Fig) for all motifs.
The performance of GAMA is position-invariant
We investigated whether the performance of our method is position-invariant, as motifs may vary in positions across biological sequences. Since our method was designed for autoregressive models, motif positions at the beginning of a sequence may be less reliably identified than those toward the end. The weaker gradients experienced for elements located towards the end of the computation graph likely contribute to this discrepancy [76]. To examine the reliability of motif identification, we implanted motifs at three different positions within the sequence: the front, middle, and end. We averaged the false-negative rate over the three motif positions (Figs 6, S14). We observed very similar false-negative-rate distributions for all three motif positions. Only slight shifts in the distributions are observable. Therefore, we concluded that the performance of GAMA is position-invariant.
The numbers on the bar charts signify the false-negative-rate on the minor y-axis, signal ratio on the minor x-axis. Major x-axis shows the results for the three motif-logic variants (AND, OR, and XOR). The major y-axis compares motifs at the front, middle, and end. They are obtained by averaging over different motif lengths at the front, middle, and end. We can see that motif recovery rates are similar for the three position conditions. We also compute the false-negative-rate for the IG baseline for comparison.
The length of the motif correlates positively with the GAMA false-negative-rate
We investigated whether longer motifs render recovery more challenging. For instance, Akbar and colleagues showed that antibody binding sites (paratopes) can vary in amino length [63, 71]. Multiple relevant elements in our sequence may dilute the IG signal. To answer this question, we queried motifs of different lengths (binary, trinary, quaternary). Binary motifs have a lower false-negative rate than the larger motifs and the false-negative-rate is similar across all three motif lengths. (S15 Fig). This can be explained by the fact that larger motifs also have more possible combinations.
GAMA false-negative-rate of retrieved motifs increases in the following order: AND-logic, OR-logic, XOR
The function of a biological sequence is driven by inter-position dependencies. Individual amino acids within binding motifs of CDR sequences can have different effects on the binding affinity. For example, a CDR might only effectively bind an antigen if all amino acids of the motif have the correct binding structure (AND-interaction). In contrast to a binding motif that just requires one of the amino acids to be correct (OR-interaction). Different interactions between amino acids within the motif can be described by logical operations. We investigated the performance of our method on different motif interactions. We considered three logic operations AND, OR, and XOR. Valid signal sequences must conform to this logic otherwise they are considered noise sequences. For example, for the AND logic to be considered valid, all signal positions must be activated, but for the OR logic only one of the sequence positions must be activated. We also included longer signals that contain more than two sequence positions; we include binary, trinary, and quaternary. The non-binary sequences are created by chaining the respective logic operation (i.e., s1 AND s2 AND s3). The false-negative-rate rate for all experiments with the same signal logic is averaged. The results are plotted in (Fig 5A “AND-logic”, Fig 5B “OR-logic”, Fig 5C “XOR-logic”). In summary, we evaluated the performance of our method across various motif interaction logics (AND, OR, XOR) and extended analyses to include longer signals, demonstrating how different logical dependencies influence false-negative rates.
GAMA attributions on synthetic antibody binding data correlate with positional contribution to binding affinities
GAMA attributions on synthetic antibody binding data correlate with positional contributions to binding affinity. We first characterized performance on a simple synthetic data set (Fig 5) and then evaluated GAMA using a more complex and realistic synthetic benchmark derived from antibody‑binding simulations generated with the Absolut! framework [44]. This framework provides binding affinity of antibody CDRH3 sequences together with the amino acid-wise binding energy values contributing to the total binding affinity [45], serving as a proxy for the position-specific ground truth amino acid significance, which is largely unattainable in experimental data. The informativeness of binding locations is likely correlated with the binding energy if a specific binding motif mediates the physical interaction with the epitope. It has to be noted that amino acids not within the paratope may also have high information content, for example, an amino acid that causes a change in the 3D structure of the paratope.
The Absolut! framework discretizes an antigen and systematically explores all possible CDRH3 binding configurations and computes the optimal binding affinity (Fig 7A). We investigated whether GAMA effectively highlights amino acid motifs that correspond to the binding energy contributions of each amino acid in the CDRH3 region.
(A) The Absolut! framework estimates CDRH3-antigen binding affinity between CDRH3 sequences and antigens by discretizing the antigen’s structural conformation and exhaustively exploring all potential binding configurations for the CDR sequence. (B) Distribution of amino acids among the top 1% binders to the 5CZV antigen. (C) The upper bar chart illustrates the average amino acid binding affinities of 4823 CDRH3 sequences, representing the top 1% binders to the 5CZV antigen as per Absolut!. The second bar chart shows the corresponding attributions assigned by GAMA, with a Spearman correlation between the CDRH3 sequences and the GAMA-importance, of 0.74. This correlation is significantly greater than zero, with a one-sided p-value of 0.004 (<0.05). The lower bar chart visualizes the IG baseline as comparison. (D) Spearman correlation analysis across four antibody datasets, each created with a different antigen, compared against the Absolut! ground truth. 95%-confidence intervals, computed by bootstrap sampling (1000 samples), are displayed as lines. Significant correlations are indicated with an asterisk (*), the significance cut-off is α = 0.05.
We generated a set of binding CDRH3 sequences for four different antigens (pdb-id 5CZV: 4823 CDRH3 sequences (Fig 7C), pdb-id 5KN5: 2047 CDRH3 sequences (S18 Fig), pdb-id 4K24: 2173 CDRH3 sequences (S19 Fig), pdb-id 3Q3G: 1013 CDRH3 sequences (S20 Fig)). We selected these CDRH3 sequences from the top 1% binders to a specific epitope on a given antigen, containing at least 1000 sequences with the lowest entropy in the dataset. (Information)-entropy quantifies the uncertainty within a dataset; the lowest entropy dataset is the one with the highest information.
For the 5CZV derived dataset (frequency distribution plotted in Fig 7B), we compared the Abolut! [44] ground-truth data with the attribution assignments derived from GAMA. Our analysis successfully identified a high binding affinity amino acid at position six while the IG baseline mostly failed to identify important features, as shown in Fig 7C. To evaluate the consistency across datasets, we calculated the Spearman correlation between Absolut! binding affinities and GAMA attributions for all four datasets, marking statistical significance with asterisks (*) in Fig 7D. The non-perfect correlation may be explained by the fact that most informative locations should not necessarily be binding locations, for example, an amino acid that causes a change in the 3D structure of the paratope. Additionally, we computed one-sided 95% confidence intervals for the correlation coefficients using bootstrap sampling, also depicted in Fig 7D. The one-sided p-values for the four antigens were as follows: 5CZV p-value = 0.004 < 0.05; 4K24 p-value = 0.018 < 0.05; 5K5N p-value = 0.08; 3Q3G p-value = 0.196. To summarize, our analyses highlight significant correlations between Absolut! binding affinities and GAMA attributions for two antigens (5CZV and 4K24), with statistical support provided by one-sided p-values and confidence intervals, while variability in other datasets underscores the complex relationship between binding locations and structural influences.
GAMA identifies the most important amino acid positions in the CDR3H region of the Trastuzumab antibody
We demonstrated that GAMA effectively identifies significant motifs in synthetic sequences and synthetic antibody sequences, even under noisy conditions. Notably, GAMA achieved best performance with datasets exhibiting less than 50% noise ratio and motifs connected by “AND”-logic. Although most other conditions surpassed the random baseline, motif identifications remained prone to errors (Fig 5). To extend our analysis, we examined the experimental Trastuzumab-HER-2 binding data from Mason et al. [72] (hereafter referred to as the “Mason dataset”). The Mason dataset comprises of 8955 sequences that were generated by mutating the original trastuzumab CDRH3 and filtering them for sequences that bind the HER-2 antigen; the resulting frequency distribution is visualized in Fig 8A.
(A) Logoplot of the 8955 sequences of the Mason dataset [72]. (B) Trastuzumab-HER-2 binding structure (pdb-ID: 1N8Z), highlighting the four binding amino acids of the Trastuzumab sequence. (C) Sequence position (x-axis) attributions according to GAMA-importance (y-axis). (D) Sequence position (x-axis) attributions according to IG baseline (y-axis). Briefly, as opposed to GAMA, the IG baseline was calculated only on the trained LSTM, without reference to a randomly initialized model. For each sequence position t, we aggregate the absolute IG values over all input token types, all output token types, and all output positions, yielding a single baseline importance score per position.
Given the experimental nature of the dataset, the amount of noise in it is unknown. Nevertheless, the crystal structure of unmutated Trastuzumab bound to HER-2 (pdb-ID: 18NZ) identifies the key paratope amino acids for binding at positions 102, 103, 104, and 105 (Fig 8B). It remains uncertain whether these positions maintain their significance within the Trastuzumab variants. Since GAMA identifies averages over all sequences, different binding motifs within the same dataset may dilute the GAMA attributions across the sequence length. Therefore, the correlation between the informativeness and the binding motif might not be perfect.
Fig 8 presents our findings with a logo plot of the amino acid frequency distribution shown in Fig 8A. The GAMA activations for the Mason dataset, highlighted in Fig 8C, underscore positions 103, 104, 105, and 107 as particularly important for the generative model, with some overlap between the predicted and known binding positions of the original Trastuzumab-HER-2 interaction. The IG baseline in Fig 8D fails to recover any of the four important positions. In summary, the analysis in Fig 8 reveals key positions identified by GAMA activations that align partially with known binding sites, emphasizing their relevance to the generative model and the Trastuzumab-HER-2 interaction.
Discussion
In this paper, we developed a novel analysis method, GAMA, to enhance the interpretation of generated sequences for autoregressive generative models and mitigate the effects of negative label bias. This advancement is particularly relevant for datasets devoid of negative label sets, known as “Positive and Unlabeled data” (PU-data), which find extensive applications beyond biological contexts, including recommender systems, fraud detection, information retrieval, and personal advertisement. Methods for PU-data typically employ semi-supervised strategies within an iterative “Two-step” learning process, wherein a supervised method classifies unlabeled data, subsequently using the new labels as training data for successive iterations [77–79]. While these PU-data approaches are versatile, they do not provide solutions specifically tailored for generative models.
Recent generative approaches to antibody sequence design often fail to elucidate why a generative model considers a newly generated sequence to be optimal for a specific task [80, 81]. GAMA addresses this gap by offering interpretability, and it is applicable to a wide range of (differentiable) autoregressive models, including large language models (e.g., protein language models), which have demonstrated the most promising recent advances in generative antibody design [82, 83], as well as other enhancements in state space models such as xLSTM [84], and Mamba [2]. GAMA could also be extended to address non-biological challenges, such as generative models for outlier detection employed in fraud detection, where the interpretability of model outputs is of importance.
Utilizing simple synthetic sequences with well-defined motif signals, we observed that GAMA can recover binding motifs under various sequence noise conditions and remains invariant to motif position (Figs 5, 6). Our results indicate a strong correlation between GAMA and the ground-truth data in simulated antibody binding data (Fig 7D). We also demonstrated the efficacy of our approach using an experimental antibody binding dataset, and thus confirming its applicability to real-world data (Fig 8).
A direct comparison to published classification based attribution studies, for example, in terms of sensitivity, is not straightforward in our setting because those approaches rely on clearly defined positive and negative labels, whereas our models are trained on 1‑class, positive-only data and there is no well-defined negative class against which to compute true‑positive or false‑positive rates. Instead, we systematically compared GAMA against an Integrated Gradients baseline computed on the trained LSTMs without reference-model subtraction; across all synthetic conditions, simulated antibody datasets, and the experimental Trastuzumab–HER2 data, GAMA consistently outperformed this IG baseline in recovering known or proxy ground-truth positions (Figs 5, 7C, 8C/8D, S3-S14, S16-S18).
The evaluation of GAMA on experimental HER-2/Trastuzumab binding data revealed that positions 103, 104, 105, and 107 are particularly significant for the generative model. When comparing these results to the four binding positions in the unmutated Trastuzumab sequence, which indicate positions 102, 103, 104, and 105 as most important for binding with the HER-2 antigen (Fig 8), we matched the positions 103, 104, 105 and 107, but failed to match position 102. This discrepancy between structural information and GAMA-identified positions could stem from several factors. Firstly, structural information is available only for the unmutated Trastuzumab sequence, making it uncertain to what extent mutated sequences retain the same binding amino acids. Secondly, GAMA provides average amino acid attributions over all sequences, meaning that multiple different motifs within the dataset could aggregate into a single “combination” motif. Thirdly, as demonstrated in Fig 5 with synthetic sequences, error increases with higher noise levels and more complex inter-position interactions.
In this study, we benchmarked GAMA using simulated paratope motifs of up to four interacting amino acids, separated by gaps of up to one or two amino acids. These motifs encompass the most probable types in terms of length and number of gaps. The distributions of paratope motifs have been thoroughly investigated [69]. For instance, the CDRH3 region exhibits a broad range of motif lengths, from one to fifteen, with a median of five, while gap numbers range from zero to five, with a median of one. In contrast, the CDRL2 region features paratope motifs ranging from one to seven in length, with a median of two, and gap numbers ranging from zero to five, with a median of two.
Applying GAMA to generative antibody sequence models represents a promising avenue for rational antibody design [25,45,85]. Understanding the rules governing antibody-antigen binding, which operate on the antibody and (protein)-antigen sequences [86], could facilitate the creation of novel, fit-for-purpose antibodies. Our work represents an initial step towards this goal.
Currently, GAMA analyzes single positions within the sequence chain, which hampers its ability to identify more complex binding patterns, such as higher-order dependency patterns [87], and may result in the erroneous identification of important positions [69]. GAMA offers an aggregated summary of the generated sequence patterns by integrating attributions across the entire dataset. This methodological approach, however, limits GAMA’s capacity to discern motifs in environments with very low signal ratios, thereby constraining its applicability to domains where motifs are anticipated to be sparse, such as Adaptive Immune Receptor Repertoire (AIRR) datasets. For example, only approximately 10% of TCR sequences in cytomegalovirus (CMV)-infected patients are associated with CMV [88].
Future work will focus on extending GAMA to single sequence analysis, enabling attributions to be computed for individual sequences. Currently, the noise of the integrated gradients for a single sequence is too high to allow for meaningful individual interpretation, necessitating the averaging of attributions over many sequences, but the gradients contain much more information in 4D, which should be further analyzed. Although the present study implements and validates GAMA using an LSTM, its computational principle does not inherently depend on the LSTM architecture: it requires a differentiable autoregressive model for which gradients can be computed with respect to the input embeddings. This suggests that GAMA could potentially be extended to other architectures, including transformers and state-space models. However, such extensions may require architecture-specific implementation and validation and remain an important direction for future work.
Additionally, future research should incorporate pairwise and higher-order dependency information, as we know that CDR binding is mediated by the interaction of multiple amino acid positions within the CDR [44,72].
Methods
Datasets
Synthetic dataset generation
Whether CDR sequences are associated with binding to an antigen usually depends on short sequences of amino acids that physically interact with the epitope (Table 1). To simulate this interaction, we created datasets containing sequences with uniform sampled amino acids, and containing implanted signal amino acids that designate a motif. Each dataset is associated with an experimental condition that is described by a 3-tuple: motif positions (list of motifs within the dataset, represented as a list of all sequence positions where the motif is implanted: (2, 4), (7, 9), (13, 15), (2, 4, 6), (7, 9, 11), (12, 14, 16), (2, 3, 4, 5), (6, 7, 8, 9), (11, 12, 13, 14)), motif logic (“AND”, “OR”, “XOR”), and noise ratio (1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1). An example of how the dataset creation is structured can be seen in Table 2. Noise ratio is the ratio between sequences with an implanted motif and noise sequences devoid of any motif if motifs are present by chance the sequence is rejected and resampled. The randomly generated sequences without signals were inserted into the datasets to assess the model’s robustness against noise. Specifically, the noise sequences were composed of uniformly sampled tokens, while the signal sequences were also uniformly sampled but included implanted signal tokens representing the motif. A token represents an element of the sequence and corresponds to an amino acid.
The synthetic datasets were generated using a deterministic seed to ensure reproducibility. Each sequence had a length of 16, a value chosen because it approximately represents the median length of human CDRH3 sequences [73]. Each experiment consisted of 10 000 signal (positive) sequences paired with a specific quantity of noise sequences. Resulting in datasets of up to 100 000 sequences for the smallest signal ratio. We demonstrated that 10 000 sequences are adequate for achieving robust model performance by training deep generative models on iteratively smaller subsets of a complete dataset (S22 Fig).
We generated new datasets based on three primary variables: signal ratio, signal positions, and signal logic. The signal position parameter also included motifs of varying sizes to examine the signal recovery rate of our model in relation to motif size.
Synthetic antibody datasets generated with the Absolut! framework
We evaluated GAMA on four antibody datasets created with the Absolut! simulation framework. Each dataset contains CDRH3 sequences that bind a single antigen. The four antigen-dataset pairs are: 5KN5 (2 047 binders), 5CZV (4 823 binders), 4K24 (2 173 binders) and 3Q3G (1 013 binders). Each dataset was generated by first identifying an epitope hotspot on an antigen, followed by selecting the top 1% of antibody binders to form the dataset. We conducted a comprehensive search of all possible epitopes within the Absolut! database and selected the four epitopes that yielded datasets with the lowest entropy (i.e., highest information content), while ensuring that each dataset contained a minimum of 1,000 sequences.
Experimental dataset
The original dataset from Mason et al. [72], comprising HER-2 binders. The positive binder sequences were filtered to include only sequences with a minimum of two reads. This filtering step was undertaken to eliminate false negatives that may have arisen due to erroneous reads or contaminations (Fig 9).
Computational methods and data processing
LSTM training and sampling
We trained an LSTM (Long Short-Term Memory) (1) model to learn a dataset in a generative manner, enabling the sampling of novel sequences. Each sequence in a dataset was encoded using one-hot encoding while adding start and stop tokens. All sequences shared the same length within one dataset. The LSTM model is trained to predict the identity function, which allows the model to predict the subsequent amino acid in the sequence, a technique known as teacher forcing [66]. This allows the sampling of novel sequences from the same distribution. We sampled new sequences from the model by initiating the trained model with a start token, then progressively sampling one amino acid at a time until an end token was reached.
The LSTM model has a hidden state size of 1024. Training was conducted using teacher forcing and mini-batches, with the ADAM optimizer [89] from PyTorch 1.8.1, utilizing default parameters except for a learning rate set to lr = 0.00001. This learning rate was determined through multiple test runs to ensure stable convergence (see S1 and S2 Figs). The weights were initialized with a uniform distribution according to where
. To ensure reproducibility of our stochastic experiments, the random seed value was set to zero. Fig 10 illustrates the prediction step of an LSTM, with input values shown in blue and output values in black, and gradients in black.
The LSTM model is queried to compute the gradient (red)IG for a given input sequence encoded with one-hot-encoding (blue)., The IG is computed from the gradient, resulting in a four-dimensional IG tensor for each sequence. The dimensions of this tensor are as follows: (1) Integrated gradients are computed for all 22 elements (comprising 20 amino acids, a start token, and a stop token) within the encoding vector for each amino acid in the sequence. (2) A sequence vector corresponding to the length of the input sequence. (3) Gradients are calculated for each dimension in the output encoding of the softmax, with the output sequence length matching the input sequence length. (4) Gradients are computed for all elements in the output sequence.
Integrated gradient analysis
Integrated Gradients (IG) [35] is a method for attributing the contribution of input features to the output of a machine learning model. Given an input, a model, and its respective output, IG quantifies the importance of each feature by measuring the degree of change in the output when that particular input feature is altered, as defined by the gradient. The importance is defined by the magnitude of change in the output resulting from a change in the feature.
Applying IG to generative models is generally non-trivial, as most attribution methods focus on explaining changes in predictions with respect to inputs, while deep generative models often lack a supervised objective that directly corresponds to the sample output. However, in the case of LSTMs, which are trained in a supervised manner to learn the identity function, IG can be computed with respect to the inputs for a single output.
IG were computed on the input sequence (encoded as a tensor with dimensions 22 x sequence length) and a zero input baseline
. The zero input baseline mirrors the dimensions of the one-hot encoded input sequence but sets all values to zero, despite this not constituting a valid one-hot encoding and does not refer to a specific token. We generated 1000 equally spaced points along a linear path between the input sequence and the input baseline, with linear interpolations performed in the one-hot encoding space.
Each input interpolation is described as for 1000 equally spaced
. For each of the 1000 input interpolations, we used the LSTM model to predict outputs and subsequently computed the gradients with respect to the corresponding inputs. By averaging over all 1000 gradients, we obtained the IG for a given sequence. These gradients were computed using automatic differentiation provided by PyTorch 1.8.1.
Calculation of integrated gradients and structuring into a 4D Tensor
The previous section described how IGs are computed. For this study, the IG are computed for each output token of a sequence, structuring the results into a four-dimensional tensor (1. Input encoding, 2. Input sequence token, 3. Output probability vector, 4. Output sequence token), which we hereafter refer to as the “IG-tensor” (refer to Fig 10). Given that IG can only be computed for a single output of the LSTM at a time, the LSTM, trained on a positively labeled dataset, was queried for the IG of an input sequence from the same dataset. This computation results in a two-dimensional tensor for a single output, representing the IG with respect to all LSTM inputs. The dimensions of this tensor are 22 (20 amino acids + start token + stop token) by the sequence length.
We performed this IG computation for all possible outputs. Each output comprises the softmax output encoding of size 21 (20 amino acids + stop token) and a prediction for each token in the input sequence of size matching the sequence length. Consequently, this results in a four-dimensional IG-tensor with dimensions [22 (number of input tokens), sequence length, 21 (number of predicted outputs), sequence length]. Following the enumeration of all possible IG computations, we reduce and summarize the 4D IG-tensor to a more manageable two-dimensional tensor in the subsequent section.
First, the 4D-IG-tensor was reduced to a 2D tensor by eliminating the first dimension (input encoding) and the third dimension (the output dimension for which the gradient is computed). The first dimension of the IG tensor was reduced by only selecting scalar values that correspond to the encoding position of the input amino acid as other positions are zero. The third dimension was reduced by averaging over all elements as we found them to be similar (S23 Fig).
The IG baseline is the mean of the absolute IG tensor across dimensions one, three, and four. This results in a one-dimensional positionwise importance score, where each sequence position t is assigned the mean absolute IG magnitude aggregated over all input token types, all output token types, and all output positions. We use this quantity as the “IG baseline” throughout the Results section. The IG baseline uses the same IG values used to compute GAMA.
GAMA computation algorithm
We summarize the computation of GAMA in Algorithm 1.
Algorithm 1. Calculation of GAMA: each sequence is passed into the trained LSTM model and the IG is computed for all outputs of the LSTM. This is done both for the reference model and the trained model. Subsequently, the difference is quantified using the GAMA function.
Packages used
For the creation of graphics, we employed several libraries. Logo plots were generated using Logomaker 0.8 [90], while all other figures were created using Seaborn 0.11.1 [91].
Our computational work relied on the following libraries: PyTorch 1.8.1 [92], NumPy 1.23.1 [93], pandas 1.2.4 [94], and SciPy 1.14.1 [95]. These libraries were integral to the implementation and analysis of our methods.
Supporting information
S1 Fig. Learning rate selection on synthetic data.
Evaluation of four different learning rates for training the LSTM on the synthetic datasets, showing that ADAM with a learning rate of 0.00001 provides a good trade-off between training speed and stability.
https://doi.org/10.1371/journal.pcbi.1014805.s001
(PDF)
S2 Fig. Learning rate selection on experimental data.
Evaluation of four different learning rates for training the LSTM on the experimental Trastuzumab–HER2 dataset, again identifying 0.00001 as a stable and effective learning rate.
https://doi.org/10.1371/journal.pcbi.1014805.s002
(PDF)
S3 Fig. IG baseline fails to recover motifs at positions 2 and 4.
Facet plot of bar charts showing IG baseline importance across sequence positions for synthetic datasets with implanted motifs at positions 2 and 4 across motif logics (AND, OR, XOR) and signal ratios; the IG baseline largely fails to highlight the true motif positions.
https://doi.org/10.1371/journal.pcbi.1014805.s003
(PDF)
S4 Fig. GAMA importance for motifs at positions 7 and 9.
Facet plot of GAMA importance across sequence positions for synthetic datasets with implanted motifs at positions 7 and 9 across motif logics and signal ratios, showing elevated importance at motif positions that diminishes gracefully with increasing noise.
https://doi.org/10.1371/journal.pcbi.1014805.s004
(PDF)
S5 Fig. IG baseline fails to recover motifs at positions 7 and 9.
Facet plot of IG baseline values for the same synthetic setups as in Supplementary Figure 4, demonstrating that the IG baseline rarely identifies positions 7 and 9 as important.
https://doi.org/10.1371/journal.pcbi.1014805.s005
(PDF)
S6 Fig. GAMA importance for motifs at positions 13 and 15.
Facet plot of GAMA importance for synthetic datasets with implanted motifs at positions 13 and 15, over motif logics and signal ratios, showing robust emphasis on the true motif positions at higher signal ratios.
https://doi.org/10.1371/journal.pcbi.1014805.s006
(PDF)
S7 Fig. IG baseline fails to recover motifs at positions 13 and 15.
Facet plot of IG‑baseline values for the same conditions as Supplementary Figure 6, indicating poor recovery of the true motif positions.
https://doi.org/10.1371/journal.pcbi.1014805.s007
(PDF)
S8 Fig. GAMA importance for motifs at positions 2, 4, and 6.
Facet plot of GAMA importance for synthetic datasets with ternary motifs at positions 2, 4, and 6 across motif logics and signal ratios, with clear emphasis on motif positions that decrease with increasing noise.
https://doi.org/10.1371/journal.pcbi.1014805.s008
(PDF)
S9 Fig. IG baseline fails to recover motifs at positions 2, 4, and 6.
Facet plot of IG baseline values for the same ternary‑motif conditions as Supplementary Figure 8, showing largely unsuccessful identification of the true motif positions.
https://doi.org/10.1371/journal.pcbi.1014805.s009
(PDF)
S10 Fig. GAMA importance for motifs at positions 2, 3, 4, and 5.
Facet plot of GAMA importance for synthetic datasets with quaternary motifs at positions 2–5 across motif logics and signal ratios, demonstrating strong importance at motif positions that is gradually reduced by noise.
https://doi.org/10.1371/journal.pcbi.1014805.s010
(PDF)
S11 Fig. IG baseline fails to recover motifs at positions 2, 3, 4, and 5.
Facet plot of IG‑baseline values for the same quaternary‑motif conditions as Supplementary Figure 10, indicating poor and inconsistent recovery of the true motif positions.
https://doi.org/10.1371/journal.pcbi.1014805.s011
(PDF)
S12 Fig. False‑negative rate of motif recovery across motif positions. Facet plot of bar charts showing false‑negative rates for GAMA (red) and the IG baseline (light red) across all 270 synthetic conditions (signal ratio, motif position, motif length, motif logic), indicating consistently lower false‑negative rates for GAMA under AND and OR logic at moderate to high signal ratios.
https://doi.org/10.1371/journal.pcbi.1014805.s012
(PDF)
S13 Fig. Top‑k error of motif recovery across motif positions. Facet plot of bar charts showing the Top‑k error metric for GAMA (red) and the IG baseline (light red) across all 270 synthetic conditions, revealing that GAMA achieves systematically better Top‑k performance than the IG baseline, particularly for AND and OR logic, and exposing improvements not visible from the false‑negative rate alone.
https://doi.org/10.1371/journal.pcbi.1014805.s013
(PDF)
S14 Fig. Top‑k position required to fully recover motifs. Bar plots of the Top‑k position metric (number of top‑ranked positions needed to recover the complete motif) for GAMA and the IG baseline across motif positions (front, middle, end) and logics (AND, OR, XOR), showing that motifs at the end of sequences are generally recovered with fewer positions.
https://doi.org/10.1371/journal.pcbi.1014805.s014
(PDF)
S15 Fig. Retrieval performance as a function of motif length.
Bar plots of false‑negative rates for motifs of different lengths and positions, demonstrating that shorter (binary) motifs exhibit lower false‑negative rates than longer (3‑ and 4‑position) motifs.
https://doi.org/10.1371/journal.pcbi.1014805.s015
(PDF)
S16 Fig. Integrated gradients on supervised LSTM binary classifiers reliably recover AND and OR motifs but fail on XOR logic.
Facet plots of mean IG attributions from supervised LSTM classifiers on synthetic motif datasets, showing strong recovery of motif positions under AND (and most OR) logic, but poor attributions under XOR.
https://doi.org/10.1371/journal.pcbi.1014805.s016
(PDF)
S17 Fig. Positive‑only IsolationForest fails to reliably classify synthetic motif sequences. Facet plots of balanced accuracy for IsolationForest across motif positions, logics, and signal ratios, indicating near‑random performance for longer motifs and only occasional moderate accuracy for shorter motifs.
https://doi.org/10.1371/journal.pcbi.1014805.s017
(PDF)
S18 Fig. Correlation between GAMA importance and binding affinities for antigen 5K5N.
Comparison of mean amino acid-level binding affinities (Absolut!) and GAMA importance scores across 2 047 top 1% CDRH3 binders to antigen 5K5N, yielding a Spearman correlation of 0.45, while the IG baseline fails to highlight key positions.
https://doi.org/10.1371/journal.pcbi.1014805.s018
(PDF)
S19 Fig. Correlation between GAMA importance and binding affinities for antigen 4K24.
Comparison of mean amino acid–level binding affinities and GAMA importance scores across 2 173 top 1% CDRH3 binders to antigen 4K24, yielding a Spearman correlation of 0.63, with poor performance of the IG baseline.
https://doi.org/10.1371/journal.pcbi.1014805.s019
(PDF)
S20 Fig. Correlation between GAMA importance and binding affinities for antigen 3Q3G.
Comparison of mean amino acid–level binding affinities and GAMA importance scores across 1 013 top 1% CDRH3 binders to antigen 3Q3G, yielding a Spearman correlation of 0.40, again with the IG baseline failing to identify informative positions.
https://doi.org/10.1371/journal.pcbi.1014805.s020
(PDF)
S21 Fig. Generative LSTM reproduces the synthetic training distribution.
Sequence logo of 10 000 samples (temperature 0.5) drawn from an LSTM trained on a synthetic dataset with an AND logic motif at positions 2 and 4 and 100% signal ratio, showing close agreement with the training frequency distribution.
https://doi.org/10.1371/journal.pcbi.1014805.s021
(PDF)
S22 Fig. Effect of dataset size on GAMA signal strength.
(A) Average GAMA importance (log scale) across sequence positions for nine synthetic datasets of varying size (1 000–9 000 sequences) with an AND logic motif at positions 2 and 4 and no noise, showing clear separation of motif and non‑motif positions even at small dataset sizes. (B) Ratio of GAMA importance at motif vs. non‑motif positions, indicating robust signal enrichment even with as few as 1 000 sequences.
https://doi.org/10.1371/journal.pcbi.1014805.s022
(PDF)
S23 Fig. Structure and reduction of IG values across encoding dimensions.
Heatmap of IG values for a single input token and next‑token prediction across all 22 encoding dimensions and output dimensions, illustrating that only the encoding row corresponding to the input token is active; this motivates reducing the 4D IG tensor to 2D by selecting the active row and averaging.
https://doi.org/10.1371/journal.pcbi.1014805.s023
(PDF)
S24 Fig. Robustness of GAMA results across random seeds.
Bar plot of GAMA importance values for a synthetic experiment with an AND logic motif at positions 2 and 4 and signal ratio 1, trained with seed = 1, showing results very similar to the corresponding panel in Fig 4 (seed = 0).
https://doi.org/10.1371/journal.pcbi.1014805.s024
(PDF)
References
- 1. Hochreiter S, Schmidhuber J. Long Short-term Memory. Neural computation. 1997 Dec 1;9:1735–80.
- 2. Gu A, Dao T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv; 2024.
- 3.
Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, et al. Generative Adversarial Nets. In: Ghahramani Z, Welling M, Cortes C, Lawrence N, Weinberger KQ, editors. Advances in Neural Information Processing Systems. Curran Associates, Inc.; 2014. https://proceedings.neurips.cc/paper_files/paper/2014/file/f033ed80deb0234979a61f95710dbe25-Paper.pdf
- 4.
Kingma DP, Welling M. Auto-encoding variational bayes. In: Bengio Y, LeCun Y, editors. 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings. 2014. http://arxiv.org/abs/1312.6114
- 5.
Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S. Deep unsupervised learning using nonequilibrium thermodynamics. In: Proceedings of the 32nd International Conference on Machine Learning. PMLR; 2015. 2256–65. https://proceedings.mlr.press/v37/sohl-dickstein15.html
- 6.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. In: Advances in neural information processing systems. Curran Associates, Inc.; 2017. https://papers.nips.cc/paper_files/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
- 7. Kucera T, Togninalli M, Meng-Papaxanthos L. Conditional generative modeling for de novo protein design with hierarchical functions. Bioinformatics. 2022;38(13):3454–61. pmid:35639661
- 8. Wu Z, Johnston KE, Arnold FH, Yang KK. Protein sequence design with deep generative models. Curr Opin Chem Biol. 2021;65:18–27. pmid:34051682
- 9. Chen Y, Wang Z, Wang L, Wang J, Li P, Cao D, et al. Deep generative model for drug design from protein target sequence. J Cheminform. 2023;15(1):38. pmid:36978179
- 10. Trinquier J, Uguzzoni G, Pagnani A, Zamponi F, Weigt M. Efficient generative modeling of protein sequences using simple autoregressive models. Nat Commun. 2021 Oct 4;12(1):1. doi:10.1038/s41467-021-25756-4
- 11. Madani A, Krause B, Greene ER, Subramanian S, Mohr BP, Holton JM, et al. Large language models generate functional protein sequences across diverse families. Nat Biotechnol. 2023;41(8):1099–106. pmid:36702895
- 12. Notin P, Rollins N, Gal Y, Sander C, Marks D. Machine learning for functional protein design. Nat Biotechnol. 2024;42(2):216–28. pmid:38361074
- 13. Hsu C, Fannjiang C, Listgarten J. Generative models for protein structures and sequences. Nat Biotechnol. 2024;42(2):196–9. pmid:38361069
- 14. Wang H, Fu T, Du Y, Gao W, Huang K, Liu Z, et al. Scientific discovery in the age of artificial intelligence. Nature. 2023;620(7972):47–60. pmid:37532811
- 15. Wu KE, Yang KK, van den Berg R, Alamdari S, Zou JY, Lu AX, et al. Protein structure generation via folding diffusion. Nat Commun. 2024;15(1):1059. pmid:38316764
- 16. Watson JL, Juergens D, Bennett NR, Trippe BL, Yim J, Eisenach HE, et al. De novo design of protein structure and function with RFdiffusion. Nature. 2023;620(7976):1089–100. pmid:37433327
- 17. Ni B, Kaplan DL, Buehler MJ. Generative design of de novo proteins based on secondary structure constraints using an attention-based diffusion model. Chem. 2023;9(7):1828–49. pmid:37614363
- 18.
Davidsen K, Olson BJ, DeWitt WS, Feng J, Harkins E, Bradley P, et al. Deep generative models for T cell receptor protein sequences. eLife. 2019 Sep 5;8:e46935. doi:10.7554/eLife.46935
- 19.
Anand N, Huang P. Generative modeling for protein structures.
- 20. Zrimec J, Fu X, Muhammad AS, Skrekas C, Jauniskis V, Speicher NK, et al. Controlling gene expression with deep generative design of regulatory DNA. Nat Commun. 2022;13(1):5099. pmid:36042233
- 21. Killoran N, Lee LJ, Delong A, Duvenaud D, Frey BJ. Generating and designing DNA with deep generative models. arXiv. 2017.
- 22. Seo E, Choi Y-N, Shin YR, Kim D, Lee JW. Design of synthetic promoters for cyanobacteria with generative deep-learning model. Nucleic Acids Res. 2023;51(13):7071–82. pmid:37246641
- 23. Riley AT, Robson JM, Ulanova A, Green AA. Generative and predictive neural networks for the design of functional RNA molecules. Nat Commun. 2025;16(1):4155. pmid:40320400
- 24. Saka K, Kakuzaki T, Metsugi S, Kashiwagi D, Yoshida K, Wada M, et al. Antibody design using LSTM based deep generative model from phage display library for affinity maturation. Sci Rep. 2021;11(1):5852. pmid:33712669
- 25. Akbar R, Robert PA, Weber CR, Widrich M, Frank R, Pavlović M, et al. In silico proof of principle of machine learning-based antibody design at unconstrained scale. MAbs. 2022;14(1):2031482. pmid:35377271
- 26.
Amimeur T, Shaver JM, Ketchem RR, Taylor JA, Clark RH, Smith J, et al. Designing feature-controlled humanoid antibody discovery libraries using generative adversarial networks. Immunology. 2020.
- 27. Shin J-E, Riesselman AJ, Kollasch AW, McMahon C, Simon E, Sander C, et al. Protein design and variant prediction using autoregressive generative models. Nat Commun. 2021;12(1):2403. pmid:33893299
- 28. Lim YW, Adler AS, Johnson DS. Predicting antibody binders and generating synthetic antibodies using deep learning. MAbs. 2022;14(1):2069075. pmid:35482911
- 29. Kong X, Huang W, Liu Y. Conditional antibody design as 3D equivariant graph translation. arXiv. 2023.
- 30. Hie BL, Shanker VR, Xu D, Bruun TUJ, Weidenbacher PA, Tang S, et al. Efficient evolution of human antibodies from general protein language models. Nat Biotechnol. 2024;42(2):275–83. pmid:37095349
- 31.
Molnar C. Interpretable machine learning. Leanpub; 2018. https://christophm.github.io/interpretable-ml-book/
- 32.
Liu Z, Qian W, Cai W, Song W, Wang W, Maharjan DT, et al. Inferring the Effects of Protein Variants on Protein–protein Interactions with Interpretable Transformer Representations. Research. 2023.
- 33.
Carter B, Krog J, Birnbaum ME, Gifford DK. Machine learning model interpretations explain T cell receptor binding. Immunology. 2023.
- 34. Goh GSW, Lapuschkin S, Weber L, Samek W, Binder A. Understanding integrated gradients with smoothtaylor for deep neural network attribution. arXiv; 2021.
- 35.
Sundararajan M, Taly A, Yan Q. Axiomatic attribution for deep networks. In: Proceedings of the 34th International Conference on Machine Learning. PMLR; 2017. 3319–28.
- 36. Ruffolo JA, Sulam J, Gray JJ. Antibody structure prediction using interpretable deep learning. Patterns (N Y). 2021;3(2):100406. pmid:35199061
- 37. Senior AW, Evans R, Jumper J, Kirkpatrick J, Sifre L, Green T, et al. Improved protein structure prediction using potentials from deep learning. Nature. 2020;577(7792):706–10. pmid:31942072
- 38.
Preuer K, Klambauer G, Rippmann F, Hochreiter S, Unterthiner T. Interpretable Deep Learning in Drug Discovery. In: Samek W, Montavon G, Vedaldi A, Hansen LK, Müller KR, editors. Explainable AI: Interpreting, Explaining and Visualizing Deep Learning. Cham: Springer International Publishing; 2019. 331–45.
- 39. Ishida S, Terayama K, Kojima R, Takasu K, Okuno Y. Prediction and interpretable visualization of retrosynthetic reactions using graph convolutional networks. J Chem Inf Model. 2019 Dec 23;59(12):5026–33.
- 40.
Nearing GS, Sampson AK, Kratzert F, Frame J. Post-processing a conceptual rainfall-runoff model with an LSTM. 2020.
- 41.
De la Fuente LA, Ehsani MR, Gupta HV, Condon LE. Towards Interpretable LSTM-based modelling of hydrological systems. EGUsphere. 2023;1–29.
- 42. Agarwal V, Reddy NJK, Anand A. Unsupervised representation learning of DNA sequences. arXiv; 2019.
- 43. Lin Y, Pan X, Shen HB. lncLocator 2.0: a cell-line-specific subcellular localization predictor for long non-coding RNAs with interpretable deep learning. Bioinform. 2021;37(16):2308–16.
- 44. Robert PA, Akbar R, Frank R, Pavlović M, Widrich M, Snapkov I, et al. Unconstrained generation of synthetic antibody-antigen structures to guide machine learning methodology for antibody specificity prediction. Nat Comput Sci. 2022;2(12):845–65. pmid:38177393
- 45.
Ursu E, Minnegalieva A, Rawat P, Chernigovskaya M, Tacutu R, Sandve GK, et al. Training data composition determines machine learning generalization and biological rule discovery. 2024
- 46.
Widrich M, Schäfl B, Pavlović M, Ramsauer H, Gruber L, Holzleitner M, et al. Modern hopfield networks and attention for immune repertoire classification. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc.; 2020. 18832–45.
- 47.
Arras L, Arjona-Medina J, Widrich M, Montavon G, Gillhofer M, Müller KR, et al. Explaining and Interpreting LSTMs. In: Samek W, Montavon G, Vedaldi A, Hansen LK, Müller KR, editors. Explainable AI: interpreting, explaining and visualizing deep learning. Cham: Springer International Publishing; 2019. 211–38.
- 48. Karpathy A, Johnson J, Fei-Fei L. Visualizing and understanding recurrent networks. arXiv; 2015.
- 49. Enguehard J. Sequential Integrated Gradients: a simple but effective method for explaining language models. In: Rogers A, Boyd-Graber J, Okazaki N, editors. Findings of the association for computational linguistics: ACL 2023. Toronto, Canada: Association for Computational Linguistics; 2023. 7555–65.
- 50. Papič A, Kononenko I, Bosnić Z. Conditional generative positive and unlabeled learning. Expert Systems with Applications. 2023;224:120046.
- 51.
Pramanik V, Maliha M, Jha SK. Enhancing integrated gradients using emphasis factors and attention for effective explainability of large language models. 2024.
- 52. Zhao Z, Shan B. ReAGent: a model-agnostic feature attribution method for generative language models. arXiv; 2024.
- 53.
Zhou W, Adel H, Schuff H, Vu NT. Explaining pre-trained language models with attribution scores: an analysis in low-resource settings. In: Proceedings of the language resources and evaluation conference. 2024. 6867–75. https://doi.org/10.63317/28zzvopwxasg
- 54. Paes LM, Wei D, Do HJ, Strobelt H, Luss R, Dhurandhar A, et al. Multi-level explanations for generative language models. arXiv; 2024.
- 55. Treppner M, Binder H, Hess M. Interpretable generative deep learning: an illustration with single cell gene expression data. Hum Genet. 2022;141(9):1481–98. pmid:34988661
- 56.
Ismail AA, Adebayo J, Bravo HC, Ra S, Cho K. Concept bottleneck generative models.
- 57.
Liu W, Li R, Zheng M, Karanam S, Wu Z, Bhanu B, et al. Towards visually explaining variational autoencoders. In: 2020 IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2020. 8639–48. https://doi.org/10.1109/cvpr42600.2020.00867
- 58.
Ross A, Chen N, Hang EZ, Glassman EL, Doshi-Velez F. Evaluating the interpretability of generative models by interactive reconstruction. In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 2021. 1–15.
- 59. Tritscher J, Krause A, Hotho A. Feature relevance XAI in anomaly detection: Reviewing approaches and challenges. Front Artif Intell. 2023;6:1099521. pmid:36844426
- 60. Sidorczuk K, Gagat P, Pietluch F, Kała J, Rafacz D, Bąkała L, et al. Benchmarks in antimicrobial peptide prediction are biased due to the selection of negative data. Brief Bioinform. 2022;23(5):bbac343. pmid:35988923
- 61. Montemurro A, Jessen LE, Nielsen M. NetTCR-2.1: lessons and guidance on how to develop models for TCR specificity predictions. Front Immunol. 2022;13:1055151. pmid:36561755
- 62. Ansari M, White AD. Learning peptide properties with positive examples only. Digit Discov. 2024;3(5):977–86. pmid:38756224
- 63. Akbar R, Robert PA, Pavlović M, Jeliazkov JR, Snapkov I, Slabodkin A, et al. A compact vocabulary of paratope-epitope interactions enables predictability of antibody-antigen binding. Cell Rep. 2021;34(11):108856. pmid:33730590
- 64. Narayanan H, Dingfelder F, Butté A, Lorenzen N, Sokolov M, Arosio P. Machine learning for biologics: opportunities for protein engineering, developability, and formulation. Trends in Pharmacol Sci. 2021;42(3):151–65. doi:10.1016/j.tips.2020.12.004
- 65.
Ramanjineyulu B, Pagadala B, Rayalu GM. Statistical inference in autoregressive models: estimation of autoregressive models. Illustrated edition. S.l.: LAP LAMBERT Academic Publishing; 2013. 260.
- 66.
Lamb AM, Alias Parth Goyal AG, Zhang Y, Zhang S, Courville AC, Bengio Y. Professor forcing: a new algorithm for training recurrent networks. In: Lee D, Sugiyama M, Luxburg U, Guyon I, Garnett R, editors. Advances in Neural Information Processing Systems [Internet]. Curran Associates, Inc.; 2016.
- 67. Laustsen AH, Greiff V, Karatt-Vellatt A, Muyldermans S, Jenkins TP. Animal immunization, in vitro display technologies, and machine learning for antibody discovery. Trends Biotechnol. 2021;39(12):1263–73. pmid:33775449
- 68. Torjesen I. The pharmaceutical journal. 2015. https://pharmaceutical-journal.com/article/feature/drug-development-the-journey-of-a-medicine-from-lab-to-shelf
- 69. Vu MH, Robert PA, Akbar R, Swiatczak B, Sandve GK, Haug DTT, et al. Linguistics-based formalization of the antibody language as a basis for antibody language models. Nat Comput Sci. 2024;4(6):412–22. pmid:38877120
- 70. Karim MR, Islam T, Shajalal M, Beyan O, Lange C, Cochez M, et al. Explainable AI for bioinformatics: methods, tools and applications. Briefings in Bioinformatics. 2023;24(5):bbad236. doi:10.1093/bib/bbad236
- 71. Scheffer L, Reber EE, Mehta BB, Pavlović M, Chernigovskaya M, Richardson E, et al. Predictability of antigen binding based on short motifs in the antibody CDRH3. Brief Bioinform. 2024;25(6):bbae537. pmid:39438077
- 72. Mason DM, Friedensohn S, Weber CR, Jordi C, Wagner B, Meng SM, et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat Biomed Eng. 2021;5(6):600–12. pmid:33859386
- 73. Shemesh O, Polak P, Lundin KEA, Sollid LM, Yaari G. Machine learning analysis of naïve b-cell receptor repertoires stratifies celiac disease patients and controls. Front Immunol. 2021;12:627813. pmid:33790900
- 74. Navin NE. Cancer genomics: one cell at a time. Genome Biol. 2014;15(8):452. pmid:25222669
- 75.
Liu FT, Ting KM, Zhou ZH. Isolation forest. In: 2008 Eighth IEEE International Conference on Data Mining. 2008. 413–22.
- 76.
Bengio Y, Frasconi P, Schmidhuber J. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies. a field guide to dynamical recurrent neural networks. 2003.
- 77. Bekker J, Davis J. Learning from positive and unlabeled data: a survey. Mach Learn. 2020 Apr 1;109(4):719–60. doi:10.1007/s10994-020-05877-5
- 78.
Zhang D, Lee WS. Learning classifiers without negative examples: A reduction approach. In: 2008 Third International Conference on Digital Information Management, 2008. 638–43. https://doi.org/10.1109/icdim.2008.4746761
- 79. Mordelet F, Vert JP. A bagging SVM to learn from positive and unlabeled examples. arXiv; 2010.
- 80. Tang X, Dai H, Knight E, Wu F, Li Y, Li T, et al. A survey of generative ai for de novo drug design: new frontiers in molecule and protein generation. arXiv. 2024.
- 81. Schneider J. Explainable Generative AI (GenXAI): a survey, conceptualization, and research agenda. arXiv; 2024.
- 82. Leem J, Mitchell LS, Farmery JHR, Barton J, Galson JD. Deciphering the language of antibodies using self-supervised learning. Patterns (N Y). 2022;3(7):100513. pmid:35845836
- 83. Dounas A, Cotet TS, Yermanos A. Learning immune receptor representations with protein language models. arXiv; 2024.
- 84. Beck M, Pöppel K, Spanring M, Auer A, Prudnikova O, Kopp M, et al. xLSTM: Extended Long Short-Term Memory. arXiv; 2024.
- 85. Wilman W, Wróbel S, Bielska W, Deszynski P, Dudzic P, Jaszczyszyn I, et al. Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discovery. Brief Bioinform. 2022;23(4):bbac267. pmid:35830864
- 86. Akbar R, Bashour H, Rawat P, Robert PA, Smorodina E, Cotet T-S, et al. Progress and challenges for the machine learning-based design of fit-for-purpose monoclonal antibodies. MAbs. 2022;14(1):2008790. pmid:35293269
- 87. Adams RM, Kinney JB, Walczak AM, Mora T. Epistasis in a fitness landscape defined by antibody-antigen binding free energy. Cell Syst. 2019;8(1):86-93.e3. pmid:30611676
- 88. Schober K, Buchholz VR, Busch DH. TCR repertoire evolution during maintenance of CMV-specific T-cell populations. Immunol Rev. 2018;283(1):113–28. pmid:29664573
- 89.
Kingma DP, Ba J. Adam: A Method for Stochastic Optimization. In: Bengio Y, LeCun Y, editors. 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings. 2015.
- 90. Tareen A, Kinney JB. Logomaker: Beautiful sequence logos in python. bioRxiv. 2019: 635029. https://www.biorxiv.org/content/10.1101/635029v1
- 91. Waskom M. seaborn: statistical data visualization. JOSS. 2021;6(60):3021.
- 92. Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, et al. PyTorch: An imperative style, high-performance deep learning library. arXiv; 2019. http://arxiv.org/abs/1912.01703
- 93. Harris CR, Millman KJ, van der Walt SJ, Gommers R, Virtanen P, Cournapeau D, et al. Array programming with NumPy. Nature. 2020;585(7825):357–62. pmid:32939066
- 94.
McKinney W. Data structures for statistical computing in python. In: Proceedings of the Python in science conference. 2010. 56–61. https://doi.org/10.25080/majora-92bf1922-00a
- 95. Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17(3):261–72. pmid:32015543