Figures
Abstract
Machine learning (ML) methods for proteins and RNAs rely on multiple sequence alignments (MSAs) and related datasets such as experimental mutagenesis libraries, yet the amount of usable information they contain remains unclear. Here, a spectral measure of information is recast into an interpretable quantity for MSAs, denoted , defined as the number of fully independent alignment positions that reproduce the observed sequence diversity. Applied to RNA MSAs, this measure shows that evolutionary constraints nearly halve diversity relative to the secondary structure alone, quantifying functional and phylogenetic restrictions beyond base pairing. The same analysis indicates even lower effective diversity in proteins, reflecting tighter packing and coevolutionary constraints.
further correlates with protein structure prediction accuracy, anticipating cases with insufficient evolutionary signal. When applied to experimentally and computationally generated libraries, it measures both produced diversity and cross-library overlap, quantifying novelty rather than redundant sampling. Together, these results establish
as an operational tool to estimate effective information in MSAs, anticipate modeling difficulties, and guide protein and RNA design.
Author summary
Machine learning has transformed biology, predicting protein structures, uncovering evolutionary rules, and designing new RNA and protein sequences. Almost every such method learns from large collections of related sequences, and the field largely assumes that more data means better models. But more is not always richer. A collection of thousands of sequences may hold far fewer independent evolutionary signals, because so many entries are near-copies shaped by shared ancestry or by designs that scarcely depart from a single template. We rarely know how much real information a dataset carries, let alone how to measure it. Here I introduce the effective length, a simple measure of the genuinely independent information in a sequence collection. It reveals that natural RNA and protein families are far more constrained than their size suggests, anticipates how reliably their structures can be predicted, and distinguishes new sequence libraries that add real information from those that merely repeat what we already know.
Citation: Opuu V (2026) A spectral framework for measuring diversity in multiple sequence alignments. PLoS Comput Biol 22(9): e1014778. https://doi.org/10.1371/journal.pcbi.1014778
Editor: Arne Elofsson, Stockholm University: Stockholms Universitet, SWEDEN
Received: February 10, 2026; Accepted: August 29, 2026; Published: September 23, 2026
Copyright: © 2026 Vaitea Opuu. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The code to reproduce all analyses is available at https://github.com/vaiteaopuu/effective_length. RNA MSAs and support-size estimates were extracted from Calvanese et al. (2024) (https://doi.org/10.5281/zenodo.10688227). Ribozyme sequences and catalytic activities were taken from Lambert et al. (2025) (https://doi.org/10.5281/zenodo.16531362). Protein MSAs were randomly sampled (n = 1000) from Mirdita et al. (2017) (https://uniclust.mmseqs.com/). MSAs used for structure-prediction benchmarks were obtained from Moussad et al. (2023) (https://doi.org/10.5281/zenodo.7682977).
Funding: The author(s) received no specific funding for this work.
Competing interests: The author has declared that no competing interests exist.
Introduction
Modern sequencing and design technologies now generate massive collections of biological sequences [1], especially for RNA [2]. These data are the primary input for machine learning (ML) models used to infer evolutionary constraints or generate functional variants. Such models support applications including mutational effect prediction from deep mutational scanning [3] and protein structure prediction from natural multiple sequence alignments (MSA), as in AlphaFold [4]. Model performance is often assumed to scale with dataset size [5], yet the amount of usable information contained in these data is usually unknown. Natural MSAs are shaped by phylogeny and uneven sampling, resulting in many closely related sequences [6,7]. Designed or experimental libraries, in contrast, densely explore narrow mutational neighborhoods around a reference [2]. Consequently, large datasets can have limited effective diversity. Estimating the information content is therefore necessary to set expectations for model performance, to assess whether new data add genuinely novel information, and to relate model capacity to the diversity present in the training data.
Several measures are commonly used to summarize sequence diversity, but their interpretation and relevance for model behavior are often unclear. Average pairwise (Hamming) distance measures how different two randomly chosen sequences are on average but can be biased by sequence clusters. Positional Shannon entropy quantifies variability at individual alignment positions [8], which is informative about local conservation but treats positions independently. Similarly, k-mer entropy summarizes the diversity of short subsequences [9], but remains insensitive to long-range constraints. The effective number of sequences downweights clusters of nearly identical sequences to correct for redundancy [10], yet depends on an arbitrary similarity threshold and yields a dataset size whose meaning for model performance is not clear [11]. As a result, these measures often yield different views of diversity because they probe only local patterns or partial summaries of variation. An ideal measure of diversity would use the MSA as a sample to estimate the size of the underlying neutral set, i.e., the sequences with sufficient fitness to perform the function. The typical approach is to model the neutral set with a distribution over sequence space, assigning to each sequence a probability of belonging to the set. The neutral set size is then estimated as the effective size of the region of sequence space covered by the distribution.
In this work, I introduce an information-based measure of sequence diversity that follows this approach: treating the MSA as a sample, it fits a probability distribution to it and estimates the effective size of the region of sequence space it covers. I express this size as an effective sequence length
through
, the number of fully independent positions that would reproduce the same size.
, is designed to capture global variation across the full alignment rather than local patterns or raw sequence counts. By construction,
depends on the collective structure of variation across sites and thus reflects correlations and constraints distributed along the sequence. Interpreting diversity on an effective-length scale allows different MSAs and experimental libraries to be compared, providing an intuitive and model-derived summary of their overall information content.
The paper is organized as follows. I first review existing measures of sequence diversity in aligned data. I then introduce the effective sequence length and describe its properties. Next, I analyze its behavior on synthetic datasets with controlled constraints to show that it captures global variation induced by correlations in the MSA. I then apply the measure to natural RNA and protein MSAs and to experimentally designed ribozyme libraries, highlighting differences in how variation is distributed across sequence space. The implementation and data are available at https://github.com/vaiteaopuu/effective_length.
Available diversity measures
Diversity and information in an MSA are not uniquely defined quantities, but rather a collection of nonequivalent summaries of variability in sequence space. Existing measures differ in the scale at which they operate (site-wise, sequence-wise, pairwise, or local contexts), in whether they rely on an explicit statistical model or are purely empirical, in their dependence on user-defined parameters, and in the type of structure they are sensitive to.
Positional entropy
Positional entropy quantifies sequence variability at the level of individual alignment positions. Let denote the empirical frequency of symbol a at position i. The positional Shannon entropy is defined as
Under an independent-site assumption, these positional entropies can be combined into a global support size,
which corresponds to the number of distinct sequences expected. These measures are effectively parameter-free aside from alphabet definition and gap handling, and are computationally inexpensive. They are widely used to assess positional conservation in MSAs and, more generally, as measures of uncertainty or support size in probabilistic sequence models, including large language models. Their main limitation is that they neglect inter-site correlations, typically leading to an overestimation of the effective information content.
k-mer diversity
k-mer diversity measures variability at the level of short, local sequence patterns. For all overlapping subsequences of length k, let denote the empirical frequency of each distinct k-mer (m) across the MSA. The Shannon entropy
defines an effective number of k-mers as . This measure is explicitly parameterized by the choice of k and is commonly used to quantify motif richness and local sequence complexity in MSAs. While easy to compute for small k, the exponential growth of the k-mer alphabet leads to sparsity and estimation issues as k increases. As with positional entropy, this measure is restricted to local patterns and does not capture long-range dependencies.
Effective number of sequences
The effective number () of sequences measures diversity correcting for redundancy in an MSA. Let
denote the fractional identity between sequences i and j, and define neighbors as sequence pairs whose similarity exceeds a fixed threshold, typically set to 80% [10]. Each sequence i is assigned a weight
where is the number of its neighbors. The effective number of sequences is then
The measure depends on the choice of similarity metric and threshold, which are typically heuristic. is widely used to mitigate sampling bias in MSAs, particularly in profile construction and coevolutionary analyses. However, it reflects only the number of non-redundant sequences under a chosen similarity criterion, rather than the diversity implied by the underlying neutral set from which the MSA is sampled.
Average pairwise distance
Average pairwise distance summarizes diversity through a fraction of mutations between sequences i and j. It is defined as
where N is the number of sequences. This quantity has a direct statistical interpretation as the expected distance between two sequences drawn uniformly at random from the alignment. Therefore, it implicitly assumes that the distribution of pairwise distances is approximately Gaussian or unimodal. Average distance is often used as a coarse measure of global divergence. Its main limitation is that it becomes uninformative when clustered or multimodal structure is present, because averaging small intra-cluster distances with large inter-cluster distances can yield a mean that does not correspond to any actual sequence relationship.
Results
Definition of 
Now I describe . Sequences of length L over an alphabet of size k are encoded as indicator vectors called one-hot (see Methods), giving
. Because each k symbol block satisfies a sum-to-one constraint, it contains only
independent degrees of freedom. To work directly in a non-redundant space, each block is projected onto a
dimensional zero-sum basis Q using Helmert’s contrastive encoding [12]. Let Z = XQ be the Helmert-encoded MSA.
Each sequence is given a weight where
which is embedded in
. Weighted centered features are
, with weighted mean
, and the covariance matrix is
.
The eigenvalues of C, , are nonnegative since C is symmetric and positive semidefinite, and satisfy
, the total variance. Normalizing as
therefore defines a distribution over orthogonal modes of variation, where each probability measures how much of the data’s variability is aligned with that direction. The spectral entropy
quantifies how broadly the observed sequence variation is distributed across the alignment. In an MSA, a high value indicates that many different patterns of variation are present across the sequences, whereas a low value indicates that the variation is concentrated in only a few recurring patterns.
I convert this spectral entropy into an effective length,
which is the effective dimensionality (or effective rank) of the MSA matrix divided by the alphabet size [13]. This quantity is interpreted as the length of a hypothetical alignment in which every position varies independently and with the full alphabet, yet produces the same total amount of variation as the observed MSA. It is not the number of positions that are actually independent; even when no position is statistically independent, quantifies the total variability as an equivalent number of independent positions. Despite the similar name, it is unrelated to the effective population size of population genetics, which counts individuals rather than alignment positions. From this, an effective support size is obtained as
representing the number of distinct sequence configurations effectively supported by the data.
The computation of is highly efficient as it scales with the feature dimension
rather than the sample size N, allowing for the processing of typical MSAs (L = 200) in a fraction of a second. In the data-poor regime (
), the framework preserves this efficiency by exploiting the duality of the eigenvalue spectrum: the covariance between sequences yields the identical set of non-zero eigenvalues as the feature covariance matrix because they share the same singular values. The approach is implemented in Python code available at https://github.com/vaiteaopuu/effective_length. For MSAs with large kL, I leverage the Singular Value Decomposition (SVD) of
so that it does not require the explicit calculation of the covariance matrix, which drastically reduces computation time, see Fig A in S1 Text.
To quantify the overlap between the sequence spaces spanned by two independent MSAs X and Y, I define the cross effective length . After weighted centering of both alignments, the cross operator is
The singular value spectrum of is processed through the same spectral–entropy pipeline used for
, yielding an effective dimensionality that measures the alignment of their principal modes of variation. A large
indicates that the two datasets explore similar regions of sequence space, whereas small values reflect weakly overlapping variability.
In conclusion, is a classic measure of information that is expressed in sequence length for interpretability. This framework is formally analogous to Principal Component Analysis (PCA) or SVD [14]. However, the method distinguishes itself technically by employing zero-sum (Helmert) encoding to resolve the one-hot limitations [2,14,15]. I show below that this encoding is critical to removing the one-hot bias. Moreover, this approach is not equivalent to simply counting how many modes explain an arbitrary variance threshold; rather, spectral entropy acts as a continuous, parameter-free summary of the effective dimensionality by weighting the contribution of the entire eigenvalue spectrum, including the tail.
Application of
to synthetic MSAs
I now examine how responds to changes in global variation under controlled conditions. To do so, I generated synthetic MSAs from a multivariate Gaussian model defined in one-hot space, where the strength of correlation across positions can be tuned directly through a single parameter
(see Methods).
I first generated MSAs of length L = 10 over an alphabet of size k = 4, using sample size N = 100 and varying from 0.1 to 1.0, see Fig 1. At
, all positions are strongly correlated in a way that only sequences with the same symbol are sampled. For each value of
, I produced 30 independent MSAs and computed four diversity measures: the effective length
, positional entropy, 3-mer entropy, the average distance, and the effective number of sequences
. To facilitate the representation in the result figure, I converted positional entropy into perplexity which is the average positional entropy exponentiated. The results are shown in Fig 2. As
increases, the model produces increasingly constrained alignments. This collapse of global variation was captured clearly by
, 3-mer entropy, and
, all of which decreased with increasing correlation. Their rates of decrease differed, reflecting their distinct sensitivities:
showed a smooth monotonic decline that tracked the reduction in global variability; k-mer entropy decreased more gradually because it captures only local motif diversity; and
decreased only once pairwise sequence identities crossed the similarity threshold (1-0.8 = 0.2) used to define redundancy. In contrast, positional entropy and average distance remained nearly constant for all
. Positional entropy remains constant because per-site symbol frequencies do not change under the chosen covariance matrix. For the average distance, the convergence into increasingly similar clusters is compensated by the higher intercluster distance (see Fig D in S1 Text). Consistently, entropy-based measures vary with MSA depth, whereas average pairwise distance remains largely insensitive (Fig G in S1 Text).
(a) Visualization of the covariance matrix for the generative model with the coupling parameter set to (fully coupled). Yellow entries indicate a correlation of 1, while other entries are 0. (b) Representative synthetic sequences generated with varying coupling strengths
, ranging from independent sites (
) to fully coordinated motifs (
). This illustrates the transition from random noise to structured patterns.
(a) Normalized diversity metrics (, 3-mer entropy, Perplexity/k, and
) plotted as a function of the coupling parameter
.
tracks global variation, which decays smoothly. (b) Scatter plot of the estimated effective support size (
) versus the true support size of the generative multivariate Gaussian model, colored by sequence length L (4 to 12). The dashed line represents the identity function (y = x).
I next assessed whether can recover the support size of the underlying multivariate Gaussian generative model. For each sequence length
and correlation value
, I approximated the true support by sampling a large number of sequences (105) and computing an empirical entropy of the resulting distribution. From much smaller samples of size N = 100, I estimated
. Across all lengths and correlations,
closely matched the ground-truth support size, lying near the identity line on a log scale, see Fig 2. I performed the same analysis using only the one-hot encoding, which resulted in a biased estimation, see Fig F in S1 Text. The latter demonstrates that the Helmert’s encoding is critical to achieve these results. Independent-site estimates derived from positional entropy consistently overestimated support in correlated conditions, whereas the estimate derived from
remained accurate even when correlations were strong. k-mer entropy is not applicable in this setting because its estimate cannot be mapped to full-length sequence support. Similarly,
does not provide a support estimate, as it reflects redundancy reduction rather than the size of an underlying sequence distribution.
Together, these results show that provides a consistent measure of global variation across sequence space, thereby offering a complementary measure of diversity. Applied to real MSAs, it should provide an accurate estimate of the diversity implied by the observed covariance structure.
Diversity of RNA MSAs
Diversity and predicted effective size of RNAs.
I now measure the diversity of real RNA MSAs, which I chose because independent support-size estimates are also available (obtained with direct coupling analysis (DCA) [15]). These MSAs are described in more detail in the Methods.
Across the 25 curated RNA MSAs, the effective length () was consistently much smaller than the alignment length (L), with an average
of 28.18 (see Fig 3), corresponding to a 4.5-fold reduction relative to the mean MSA length (
). This indicates strong evolutionary and structural constraints. Notably, this value is substantially lower than what is expected from secondary-structure constraints alone. For instance, generating 500 sequences constrained to the tRNA fold using RNAinverse [16] yielded
= 37.37, whereas the natural tRNA MSA (
sequences) extracted from [15] produced
. RNAinverse is a method that generates sequences predicted to adopt a specified target secondary structure. Extending this comparison to all 25 families, sequences generated solely from their consensus secondary structures (computed with ViennaRNA [16]) yielded an average
, still 2.23-fold below the alignment length but approximately twice the diversity observed in natural MSAs.
(a) Density histogram showing the distribution of the normalized effective length () for the 25 natural RNA MSAs (blue) and DCA generated libraries (orange). (b) Scatter plot of the spectral support size estimate (
) versus the independent estimate (
) on a
scale. Grey dots indicate support sizes calculated using the full MSA. Blue dots indicate
support sizes calculated by reweighting sequences with the the redundancy reduction derived from
[10]. Orange dots indicate the
with redundancy removed and regularization. (c) Diversity saturates as the size of MSA increases. For each RNA MSA, I subsample a fraction of the sequences and computed
. I then evaluated the marginal gain in diversity as a function of sample size, quantified by
. The curve shows the average across all RNA families (solid line), and the shaded region indicates the 95% confidence interval.
In contrast, generative models populate markedly different volumes of sequence space. Libraries sampled from DCA models for the same families exhibited a much higher average , suggesting that such models can explore a sequence space approximately 2.48 times larger than natural diversity, see Fig 3. By comparison, sequences generated with a variational autoencoder (VAE) yielded
, closely matching the natural MSAs. These results indicate that models trained on the same data can produce substantially different effective sequence spaces. However, higher diversity does not necessarily imply functionality; this aspect is examined in the following section on experimental ribozyme libraries.
I further assessed whether provides a meaningful estimate of support size by comparing
with independent support-size estimates obtained from a DCA variant (eaDCA) [15] across the same 25 RNA families. Using full MSAs, the log–log relationship between the two measures yielded a Pearson correlation of
(Fig 3). Applying the classical redundancy reweighting implemented in [10] increased the correlation to
, and adding regularization (
) further improved it to
.
Finally, I observed a connection between and the RNA folding thermodynamics. This diversity measure corresponds to the length of a fully random MSA, which can be understood as the alignment of an RNA that is completely unfolded. In such a state, the molecule contributes only configurational entropy, which is proportional to its length. I tested this interpretation using the MSAs and comparing
with the folding free energy of the consensus secondary structure predicted by RNAalifold [16], obtaining a Spearman correlation of
(p-value = 0.003, shown in Fig 4). Using sequences generated using RNAinverse on randomly generated folds of exactly the same length but varying stability (controlled by the number of paired nucleobases), I obtained an even higher correlation of
(see Fig E in S1 Text). These results provide evidence of the proposed connection.
Scatter plot showing the relationship between the consensus folding free energy ( in kcal/mol) on the y-axis and the normalized effective length (
) on the x-axis for 25 RNA families. The plot reports the Spearman rank correlation (
) and the associated p-value (0.003).
Taken together, these results show that natural RNA families explore only a fraction of the sequence space compatible with their secondary structure, while generative models can access broader regions of this space. The observed association with folding stability further establishes the interpretability of .
Diversity of experimental library of RNA sequences.
I now measure the diversity of experimentally generated RNA sequence libraries by analyzing the comprehensive dataset produced in [2]. In this study, multiple generative modeling approaches were used to design large sets of the group I intron of the Azoarcus bacterium, which were subsequently synthesized and assayed for catalytic activity. Here, I analyze the diversity of the active sequences only unless otherwise mentioned (5895 sequences). The resulting collection provides a unique opportunity to evaluate how different computational design pipelines populate sequence space and how their diversity changes as variants accumulate mutations away from the wild type. As high-throughput experimental screens of model-generated sequences are becoming increasingly common, establishing a principled way to quantify and compare the diversity of such libraries is essential. Here, I omit k-mer entropy, whose interpretation is unclear, and Neff, which in this regime either remains close to one when normalized or simply reflects the number of sequences in each bin. All results are shown in Fig 5.
Diversity metrics plotted against the number of mutations from the wild type (WT) sequence for various generative models. (a) Average pairwise Hamming distance between sequences. (b) The effective sequence length (). (c) Average positional perplexity. The different colored lines correspond to the specific generative models and baselines (e.g., DCA, VAE, RDM, CHI) listed in the legend [2]. (d) Detecting mode collapse and coverage. Scatter plot of the normalized cross effective length
against the average nearest-neighbor similarity NN. Each point is a sample (N = 4000) from a DCA model trained on family RF01734, annotated by sampling temperature.
Average pairwise distance. The distance grows approximately linearly with the number of mutations from the wild type, indicating that as mutations accumulate, the generated functional sequences become increasingly dissimilar. RDM reaches the largest values among methods, as it samples mutations fully at random (uniformly across positions and nucleotides). However, this result is misleading because it does not account for the very small number of active RNA sequences found in the RDM dataset. The average pairwise distance is sensitive to clustered structure and implicitly assumes a unimodal distance distribution: a small number of well-separated functional clusters can yield a large average even if within-cluster diversity is low. Moreover, removing sequences that interpolate between clusters can increase the mean, while adding a distant sequence that forms a new cluster may leave it nearly unchanged, even though the underlying structural diversity has increased. Consequently, this measure provides a misleading picture of diversity because the mean distance is not monotonic with the true geometric organization of the dataset.
Positional entropy. To make this measure more interpretable, I converted positional entropy to perplexity, given by the exponential of the average site-wise entropy. This quantity reaches its maximum at approximately 50 mutations in the CHI dataset, which is derived from natural homologs with indels replaced by the wild-type nucleotide. This result is inconsistent with the average pairwise distance. In structured RNAs, however, sequence variation is strongly constrained by conserved stems, loops, and long-range base-pairing interactions. Because positional entropy treats sites independently, it necessarily ignores these structural constraints and therefore counts variability at positions that cannot vary freely in combination.
Effective sequence length .
captures global diversity and therefore separates model behaviours more clearly. Most approaches exhibit a steady decrease in
as mutation counts rise, which is consistent with the fact that only a narrow subset of highly mutated variants remains functional as mutations accumulate. DCA achieves the highest
over intermediate mutation ranges. At larger mutational distances (
), the hybrid DCA–SB model maintains higher values. The VAE, despite recovering functional sequences at rates comparable to DCA, produces noticeably lower diversity. This result is consistent with the reduced diversity observed in the above section for VAE.
Identifying mode collapse when generating new samples is a critical sanity check. To do so, I use two complementary measures: the overlap given by the cross effective length and the average similarity to the training given by the average nearest similarity. The normalized cross effective length
tells whether a generated library X covers the full training MSA () or only a subset of its modes: a value below 1 means only part of the training variability is recovered, the signature of collapse. Coverage alone is not sufficient, however, because a library can span all the training modes (
) while lying far from the training sequences. I complement it with the nearest-neighbor similarity NN to the training set, which measures how close generated sequences actually stay to the training sequences (see the definition in SI). Together, the two characterize the regimes: low coverage identifies collapse, while full coverage combined with high nearest-neighbor similarity identifies faithful, non-collapsed generation that remains near the functional data. I note that diversity-based measures alone cannot certify that a low
is genuine constraint rather than mode collapse without a ground-truth functional space; the measured fraction of active sequences is what resolves this. Fig 5 illustrates this: I trained a DCA model on an RNA family and generated samples at temperatures ranging from T = 0.03 to T = 3. The results show collapse at low temperature, increasing to broad, noisy overlap at T = 3.
On the ribozyme libraries (including active and inactive sequences), coverage is the main discriminator. DCA at T = 1 covers the CHI reference distribution fully (), whereas T = 0.3 drops to a subset (
) despite producing more sequences, the collapse signature, at comparable similarity (
). Random mutagenesis shows the complementary case: broad coverage (
) with somewhat lower similarity (
), spanning the same modes while sampling slightly farther from the functional sequences. I chose to include both actives and inactives to avoid mixing the generation biases with the activity selection pressures.
Average pairwise distance systematically overestimates diversity because a few distant clusters inflate the mean despite low within-cluster variability. Positional entropy is more stable but ignores long-range interactions, which inflates diversity when coordinated constraints are present. , by incorporating global correlations, provides a more faithful estimate of the effective functional diversity explored by each design strategy. Moreover, the combination of
and NN allows one to detect mode collapse and off-manifold exploration.
Diversity of protein MSAs for fold prediction
I next test if the MSA diversity is related to the protein structure prediction performance with four state-of-the-art deep learning models: AlphaFold [4], RoseTTAFold [17], ESMFold [18], and OmegaFold [19]. This analysis utilized protein MSAs from the study [20], focusing on the correlation between the diversity and two key quality metrics: the Local Distance Difference Test (LDDT) [4] and the Template Modeling Score (TM-score) [21].
Across the 60 analyzed protein families, the average is , which is lower than the diversity observed in RNA. In Fig C in S1 Text, I vary the alphabet size while keeping the sequence length and correlation structure fixed, showing that
is comparable across alphabet sizes. To further support the proteins lower diversity, I computed
across 1000 MSAs taken randomly from the Uniclust30 database [22] which contains one MSA per cluster of 30% identity defined in the UniProt database. The obtained average
. This result suggests that the evolutionary constraints applied on proteins are typically stronger than for RNAs.
To quantify how much of this reduction comes from inter-position coupling, I contrast with the effective length computed from positional entropy
alone,
, which ignores correlations. The ratio
is 2.02 for RNA and 6.27 for proteins, indicating that coupling compresses effective diversity about three times more strongly than in RNA. This is consistent with coevolutionary analyses in which RNA families are captured by a few couplings tied to secondary-structure base pairs [23], whereas proteins require a denser coupling network.
While the sequences used to train structure prediction models are essentially the same, the ways in which they are used vary drastically: in AlphaFold and RoseTTAFold, the model input is an MSA; in ESMFold and OmegaFold, the input is a single sequence. Despite these differences, a positive correlation is observed between and model TM-scores across all platforms except AlphaFold (see Fig B in S1 Text and Table 1). In contrast, no significant correlation was found between the number of sequences N per MSA (or sequence length L) and prediction accuracy. Similarly,
did not show any significant correlation either, which confirms the results published earlier in [11]. For LDDT, correlations were stronger for all models except AlphaFold.
displayed significant but weak associations. These results suggest that
could help in anticipating how challenging a prediction is.
The association strength with MSA diversity varies across the two types of models. For LDDT, single-sequence models exhibited the strongest associations, with Spearman correlations of for OmegaFold and
for ESMFold. I chose Spearman because there is no reason to believe that the relation between performance and diversity is linear. RoseTTAFold followed with
, while AlphaFold showed a more moderate correlation of
. The weaker correlation observed for AlphaFold likely reflects its highly compressed dynamic range: over 80% of AlphaFold predictions on CASP15 exceed LDDT = 0.72, leaving little variance for
to explain, whereas single-sequence models such as OmegaFold span a broader performance range. The global fold accuracy, measured by TM-score, followed a similar trend; ESMFold and OmegaFold demonstrated the most robust connection to sequence diversity (
and
respectively), whereas AlphaFold showed no significant correlation (p-value = 0.23). These benchmarks were performed on the CASP15 dataset [20]. Although the amount of information available is similar, their performances from single-sequence-based to MSA-based models vary notably.
To assess whether a protein family contains sufficient evolutionary information to support accurate structure prediction, I estimated model-specific thresholds based on LDDT. I define high local accuracy as LDDT > 0.7, consistent with commonly used confidence interpretations in [4]. For each predictor, I determined a cutoff
such that families with
satisfy
. Although some architectures operate on single sequences, an auxiliary MSA was still constructed solely to quantify family diversity; the alignment itself was not used as model input.
therefore acts as a diagnostic filter: computed from the alignment alone, at a negligible cost compared to the prediction itself, it flags families whose evolutionary signal is too weak for accurate modeling before any prediction is run. The resulting thresholds are: ESMFold (
), AlphaFold (
), OmegaFold (
), and RoseTTAFold (
).
Prediction accuracy is associated by . The strong correlation between
and the performance of single-sequence models, together with the weaker dependence observed for MSA-based models, suggests that alignment procedures further enhance the evolutionary signal used for protein structure prediction. These results provide an operational threshold of diversity to estimate the accuracy in structure predictions.
Discussion
This work introduces an information-theoretic framework to quantify sequence diversity in multiple sequence alignments through the effective sequence length, . In addition to its interpretability in sequence length, it differs from PCA-based measures used so far because it relies on a zero-sum projection that is critical for the accuracy of it. By leveraging spectral entropy,
provides a measure of global variation that captures correlations distributed across positions, rather than local or pairwise summaries. Synthetic benchmarks demonstrated that
reliably tracks reductions in accessible sequence space induced by increasing constraints, whereas commonly used measures such as positional entropy or average pairwise distance fail to do so under correlated constraints.
Applied to natural RNA families, reveals that RNA MSAs occupy a highly restricted region of sequence space, with an average effective length of 28.18 positions, approximately 4.5-fold smaller than the alignment length. This constraint is substantially stronger than what is expected from secondary structure alone, as shown by comparisons with sequences generated to satisfy identical folds. This indicates that evolutionary selection imposes strong constraints beyond base pairing, likely reflecting requirements on folding kinetics, tertiary contacts, and functional robustness. The observed correlation between
and folding free energy further supports a physical interpretation of
as a proxy for configurational entropy, linking sequence diversity to thermodynamic stability.
Proteins exhibit even smaller effective lengths than RNAs, despite their larger alphabet, indicating even stronger evolutionary and structural constraints. Mechanistically, RNA folding is driven mainly by base pairing, with tertiary contacts contributing comparatively little to stability [24], so its variability is dominated by a sparse, largely pairwise coupling structure. Folded proteins instead fold through a denser network of hydrophobic-core packing and tertiary contacts, captured by the coevolutionary couplings that coincide with native contacts [10] and reproduce collective variability missed by independent-site models [25]. Part of the RNA–protein gap may also be methodological: RNA MSAs are typically built with structure-aware covariance models [26], which imprint the secondary-structure signal into the alignment [27].
Across protein families, structure prediction accuracy correlates with but not with the raw number of sequences in the MSA. This result shows that model performance scales with informative diversity rather than dataset size or effective sequence count. The particularly strong dependence observed for single-sequence models suggests that MSA-based methods implicitly amplify evolutionary information, partially compensating for limited intrinsic diversity. In contrast, when such amplification is absent, prediction accuracy becomes directly constrained by the effective richness of the underlying sequence family. The association between
and prediction accuracy is not necessarily causal. Accuracy is measured as the distance (LDDT or TMscore) to a single deposited reference structure, so a low value can arise either because the MSA lacks informative evolutionary diversity (low
) or because the target is not described by a single conformation, as in fold-switching or allosteric systems. The latter is a property of the single-structure ground truth rather than of
, since only one conformation is available per target and targets without a stable fold are excluded by construction. Building on these observations, I propose empirical diversity thresholds to estimate the minimum
required to reliably achieve high structural prediction accuracy. These thresholds turn
into a screening step ahead of structure prediction, identifying the families for which additional sequences, rather than more computation, are the limiting factor.
The analysis of experimentally generated ribozyme libraries illustrates the utility of for evaluating generative models, particularly in light of the limitations of commonly used measures such as average pairwise distance and positional entropy, which respectively overestimate diversity through cluster separation and ignore long-range constraints. Applied to the libraries, covariance-based models explore a broader functional sequence space than variational autoencoders, even when both achieve comparable fractions of active sequences. Consequently,
can serve as an operational tool to guide sequence design toward underexplored regions of sequence space, for example within active learning frameworks.
measures the diversity implied by the second-order statistics of an MSA. The covariance matrix summarizes the observed sequence distribution through the joint variation of pairs of positions; therefore, constraints involving three or more positions at once enter only through their pairwise approximation. Enzymes are the clearest case where this matters, as a few reactive residues must be held in a precise geometry to stabilize its transition state. This arrangement couples the catalytic positions and the scaffold supporting them collectively rather than pair by pair. However, the decomposition is not just a collection of independent pairs: the eigenvectors of C spread over many positions and recover the collective modes known as sectors [28], so a network of pairwise terms produces multi-site groups of covarying positions. This is the level of description on which coevolutionary models operate, including the mean-field DCA that inverts the covariance matrix [10]. Such a model type is actually sufficient in practice to design catalytically active enzymes [29] and ribozymes [2]. Extending the spectral pipeline to higher-order statistics would capture the residual epistasis directly, at a sampling cost that grows quickly with the order. Constraints invisible to C make an alignment appear more variable than it is, so
is an upper bound on effective diversity, which is the conservative direction for its use as a diagnostic.
To complement , I introduced a cross effective length
, a model-based measure of how much one sample covers the variability of another, paired with a nearest-neighbor similarity to the training set. Together they provide a sanity check for mode collapse: low coverage of the training distribution signals collapse, whereas broad coverage at low similarity signals off-manifold exploration. This diagnostic operates relative to the training data; without a ground-truth functional space, it can flag collapse but cannot certify that a low diversity reflects genuine biological constraint.
More broadly, expressing diversity on an effective-length scale enables quantitative decision rules for sequence datasets. can be used to set minimum diversity thresholds before training predictive models, to stop data collection once additional sequences no longer increase effective information, and to select new designs that maximize diversity gain rather than raw count. In iterative or active-learning pipelines, it can serve directly as an objective function to bias sampling toward underrepresented regions of sequence space. Because it is model-agnostic and comparable across datasets, the same criterion can guide curation, experimental allocation, and cross-library comparison whenever correlations dominate variability. As generative models proliferate,
offers a sanity check for the AI era: it quantifies whether a new library adds genuine information or merely more samples of the same constraints.
Methods
One-hot
Sequences are represented using a standard one-hot encoding scheme to map discrete symbols into a vector space. A sequence of length L over an alphabet of size k (e.g., k = 5 for RNA, including the gap) is encoded as a binary vector . For each position
, the state is represented by a block of k binary variables, where the entry corresponding to the observed symbol is set to 1 and all others to 0. An MSA containing N sequences is therefore represented as a binary matrix
, where each row corresponds to a single encoded sequence.
Helmert encoding of MSAs
Helmert encoding replaces categorical one-hot vectors with an orthonormal contrast basis that removes the linear dependence inherent in one-hot representations [12]. As an example, I consider here RNA sequences with four-nucleotide alphabet {A,C,G,U}. The encoding uses the orthonormal contrast matrix:
The columns of Q span the zero-sum subspace and are orthonormal, yielding successive contrasts: first comparing A with C, then the mean of {A,C} with G, and finally the mean of {A,C,G} with U. Multiplying each one-hot nucleotide vector by Q maps it to a three-dimensional orthonormal coordinate, preserving all information except the redundant overall offset. Applying this transformation along an RNA sequence produces a compact and statistically well-conditioned representation suitable for downstream modeling. This procedure is generalizable to amino acids as well. For MSAs, I include gaps as an additional symbol.
Synthetic MSA generation
To produce synthetic MSAs with controlled amounts of coordinated variation, we construct a covariance directly in one-hot space and tune a coupling parameter . At each position, I begin with the standard categorical covariance, which captures how symbols fluctuate relative to one another at a single site. To introduce dependence across positions, I mix this position-wise structure with a fully coupled component. In this coupled term, a given symbol (e.g., “A”) at one position co-varies only with the same symbol at all other positions and is anticorrelated with the other symbols. As
increases, the model increasingly favors configurations in which many positions simultaneously shift toward the same symbol. When
is large, the resulting alignments are dominated by sequences that are composed mostly of a single symbol pattern, whereas small
yields nearly independent site variation. Gaussian samples drawn from this covariance are decoded blockwise by selecting the largest component, producing categorical sequences that reflect the intended level of global coordination. The specific algorithm to sample sequences from the covariance matrix is described in supplementary information.
Phylogenetic bias in MSAs
An MSA is not a sample of independent draws from the neutral set: sequences are related by a phylogeny, so they share variation inherited from common ancestors. Sampling is uneven on top of this, since databases follow sequencing effort rather than evolutionary diversity. The extreme case arises in pandemic surveillance, where a single bacterial or viral species is sequenced tens of thousands of times over the course of an outbreak, and the resulting alignment is dominated by a recently expanded clade whose members differ by a handful of positions. The variation of such an MSA is then carried by the few sites that separate the oversampled clade, the eigenvalue spectrum of C concentrates on a small number of modes, and decreases accordingly. This reduction reflects the composition of the sample rather than a genuine loss of variability in the underlying neutral set. In such cases, the uniform weights
are replaced by the redundancy weights described in the Effective number of sequences section,
, where
is the number of sequences within 80% identity of sequence i, so that each cluster of similar sequences is reduced to a single effective observation. This is the reweighting applied in Fig 3, where it raises the agreement with the independent support size estimates from
to
.
Natural RNA MSAs
To estimate available diversity in natural MSAs, I extracted RNA families from one study. I incorporated the 25 RNA families from [15] for which independent estimates of effective support size were previously obtained. The average MSA size is 3799 sequences. This dataset enables a direct comparison between and established model-based estimates of variability. Together, these two sources provide a broad and structurally grounded benchmark for quantifying how much meaningful variation natural RNA families contain.
I constructed an experimental benchmark using data from a high-throughput study in which multiple generative modeling approaches were used to design variants of the Azoarcus ribozyme (197 nucleotides) and experimentally measure their activity. The study evaluated a broad spectrum of sequence-generation methods, including simple baselines—random uniform mutagenesis (RDM) and independent profile sampling (PRO)—structure-aware mutational schemes (BPR, BPR-3D, SB, SB-3D), evolution-based statistical models derived from natural intron alignments (DCA sampled at T = 1 and T = 0.3), a variational autoencoder trained on the same data (VAE), a hybrid Potts–structure approach (DCA-SB), and a reference set of chimeric natural introns (CHI). For each of these design strategies, I extracted all sequences that were synthesized and experimentally assayed together with their measured activities at . After applying the quality and filtering criteria used in the original screen, this yielded a comprehensive dataset of designed ribozyme variants paired with quantitative activity measurements, which I use to evaluate how diversity measures relate to only functional sequences. The dataset is composed of 15146 entries where 5895 were found active (39%).
Supporting information
S1 Text. Supplementary figures. Contains Figs A–G supporting the main text.
Fig A. Runtime of the effective length computation as a function of sequence length L for fixed sample size N = 100 and alphabet size k = 4. Synthetic MSAs are generated from a Gaussian covariance model with correlation parameter . Execution time is reported for two implementations: singular value decomposition and covariance eigen-decomposition. Each point corresponds to the mean wall–clock time over independent realizations. Fig B. Relationship between sequence diversity (
) and protein structure prediction quality. (a) Scatter plots showing the correlation between the effective length (
) and the Local Distance Difference Test (LDDT) for four transformer-based models: AlphaFold2, RoseTTAFold, ESMFold, and OmegaFold. Spearman correlation coefficients (
) and p-values are shown for each model, with single-sequence PLM-based models (ESMFold and OmegaFold) exhibiting the strongest dependencies on family diversity. (b) Corresponding correlations for the global TM-score. All benchmarks were performed on the CASP15 dataset [20]. Fig C. Alphabet size effect on
. Synthetic MSAs of size N = 500 and
but alphabet size k is varying from 4 to 20. The
is recorded and shown in blue. Fig D. Synthetic MSA with varying correlation strength between positions and symbols. In the first row (a), I show the MSA, where one colour corresponds to a nucleotide and each row is a sequence. In the second row (b), I show the PCA projection of the MSA, showing how they converge to four typical sequences. Fig E. Folding
correlates with
. I generated 500 variants using RNAinverse on 40 randomly generated structures. I computed
for each of the 40 datasets together with the average folding stability (computed with ViennaRNA [16]), which is shown as a scatter plot. Fig F. Scatter plot of the estimated effective support size (
) using directly the one hot encoding instead of the Helmert encoding, colored by sequence length L (4 to 10). The dashed line represents the identity function (y = x). Fig G. Effect of MSA depth on diversity estimates. Synthetic MSAs generated with fixed correlation (
) are analyzed as a function of sample size N. Entropy-based measures (positional entropy and k-mer entropy) vary with depth, whereas average pairwise distance remains largely insensitive.
https://doi.org/10.1371/journal.pcbi.1014778.s001
(PDF)
References
- 1. Notin P, Kollasch A, Ritter D, Van Niekerk L, Paul S, Spinner H, et al. Proteingym: Large-scale benchmarks for protein fitness prediction and design. Advances in Neural Information Processing Systems. 2023;36:64331–79.
- 2. Lambert CN, Opuu V, Calvanese F, Pavlinova P, Zamponi F, Hayden EJ, et al. Exploring the space of self-reproducing ribozymes using generative models. Nat Commun. 2025;16(1):7836. pmid:40846705
- 3. Meier J, Rao R, Verkuil R, Liu J, Sercu T, Rives A. Language models enable zero-shot prediction of the effects of mutations on protein function. Advances in Neural Information Processing Systems. 2021;34:29287–303.
- 4. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–9. pmid:34265844
- 5. Rives A, Meier J, Sercu T, Goyal S, Lin Z, Liu J, et al. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc Natl Acad Sci U S A. 2021;118(15):e2016239118. pmid:33876751
- 6.
Ding F, Steinhardt J. Protein language models are biased by unequal sequence sampling across the tree of life. In: ICLR 2024 Workshop on Generative and Experimental Perspectives for Biomolecular Design.
- 7. Lupo U, Sgarbossa D, Bitbol A-F. Protein language models trained on multiple sequence alignments learn phylogenetic relationships. Nat Commun. 2022;13(1):6298. pmid:36273003
- 8. Schneider TD, Stephens RM. Sequence logos: a new way to display consensus sequences. Nucleic Acids Res. 1990;18(20):6097–100. pmid:2172928
- 9. Vinga S, Almeida J. Alignment-free sequence comparison-a review. Bioinformatics. 2003;19(4):513–23. pmid:12611807
- 10. Morcos F, Pagnani A, Lunt B, Bertolino A, Marks DS, Sander C, et al. Direct-coupling analysis of residue coevolution captures native contacts across many protein families. Proc Natl Acad Sci U S A. 2011;108(49):E1293-301. pmid:22106262
- 11. Guo Z, Wu T, Liu J, Hou J, Cheng J. Improving deep learning-based protein distance prediction in CASP14. Bioinformatics. 2021;37(19):3190–6. pmid:33961009
- 12.
Fox J, Weisberg S. An R companion to applied regression. Sage Publications. 2018.
- 13.
Roy O, Vetterli M. The effective rank: A measure of effective dimensionality. In: 2007 15th European Signal Processing Conference, IEEE. 2007. 606–10.
- 14. Baxter-Koenigs AR, El Nesr G, Barrick D. Singular value decomposition of protein sequences as a method to visualize sequence and residue space. Protein Sci. 2022;31(10):e4422. pmid:36173173
- 15. Calvanese F, Lambert CN, Nghe P, Zamponi F, Weigt M. Towards parsimonious generative modeling of RNA families. Nucleic Acids Res. 2024;52(10):5465–77. pmid:38661206
- 16. Lorenz R, Bernhart SH, Höner zu Siederdissen C, Tafer H, Flamm C, Stadler PF, et al. Viennarna package 2.0. Algorithms for molecular biology. 2011;6(1):26.
- 17. Baek M, DiMaio F, Anishchenko I, Dauparas J, Ovchinnikov S, Lee GR, et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 2021;373(6557):871–6. pmid:34282049
- 18. Lin Z, Akin H, Rao R, Hie B, Zhu Z, Lu W, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123–30. pmid:36927031
- 19. Wu R, Ding F, Wang R, Shen R, Zhang X, Luo S, et al. High-resolution de novo structure prediction from primary sequence. BioRxiv. 2022;:2022–07.
- 20. Moussad B, Roche R, Bhattacharya D. The transformative power of transformers in protein structure prediction. Proc Natl Acad Sci U S A. 2023;120(32):e2303499120. pmid:37523536
- 21. Zhang Y, Skolnick J. Scoring function for automated assessment of protein structure template quality. Proteins. 2004;57(4):702–10. pmid:15476259
- 22. Mirdita M, von den Driesch L, Galiez C, Martin MJ, Söding J, Steinegger M. Uniclust databases of clustered and deeply annotated protein sequences and alignments. Nucleic Acids Res. 2017;45(D1):D170–6. pmid:27899574
- 23. De Leonardis E, Lutz B, Ratz S, Cocco S, Monasson R, Schug A, et al. Direct-Coupling Analysis of nucleotide coevolution facilitates RNA secondary and tertiary structure prediction. Nucleic Acids Res. 2015;43(21):10444–55. pmid:26420827
- 24. Tinoco I Jr, Bustamante C. How RNA folds. J Mol Biol. 1999;293(2):271–81. pmid:10550208
- 25. Figliuzzi M, Barrat-Charlaix P, Weigt M. How Pairwise Coevolutionary Models Capture the Collective Residue Variability in Proteins?. Mol Biol Evol. 2018;35(4):1018–27. pmid:29351669
- 26. Nawrocki EP, Eddy SR. Infernal 1.1: 100-fold faster rna homology searches. Bioinformatics. 2013;29(22):2933–5.
- 27. Opuu V. Hybrid generative model: Bridging machine learning and biophysics to expand rna functional diversity. bioRxiv. 2025;2025:2025–01.
- 28. Halabi N, Rivoire O, Leibler S, Ranganathan R. Protein sectors: evolutionary units of three-dimensional structure. Cell. 2009;138(4):774–86. pmid:19703402
- 29. Russ WP, Figliuzzi M, Stocker C, Barrat-Charlaix P, Socolich M, Kast P, et al. An evolution-based model for designing chorismate mutase enzymes. Science. 2020;369(6502):440–5. pmid:32703877