This is an uncorrected proof.
Figures
Abstract
Archaea represent one of the three domains of cellular life and yet account for fewer than 1% of experimentally determined protein structures, leaving the extent of their structural novelty unknown. Here we present a systematic domain-level classification of 124,075 proteins from 65 archaeal classes spanning 21 phyla and all major lineages, using both AFDB and newly predicted AlphaFold3 structures classified against the Evolutionary Classification of protein Domains (ECOD). Archaeal proteins span 987 of the 2,457 ECOD X-groups defined across all cellular life, roughly 40% of known fold diversity captured within a single domain of life. Clustering by Foldseek recovered structural relationships for 63% of domains that are singletons by sequence comparison. To characterize the 21% of proteins lacking high-confidence classification, we applied successive filters for structure prediction confidence, protein length, and structural cluster context, reducing 8,452 domain-free proteins to a small number of well-folded structural orphans (less than 0.1% of the dataset). The unclassified fraction is dominated by sub-threshold matches (matches below the 0.85 DPAM confidence cutoff for high-confidence T-group assignment) to known folds (14% of all proteins) and low-confidence structure predictions (5%), not by novel structures. These results demonstrate that the protein fold repertoire at the single-domain level is broadly conserved across the deepest phylogenetic distances in cellular life, and that the gap between archaeal and well-characterized proteomes reflects classification sensitivity for divergent sequences rather than unexplored structural diversity.
Author summary
Archaea are the least-studied of the three domains of life, yet they matter enormously: they flourish in extreme environments, drive global nutrient cycles, and are closely tied to the origin of eukaryotes, the group that includes our own cells. What most of their proteins do, and where those proteins came from, remains largely uncharted. Proteins are built from folded modules called domains. A domain’s three-dimensional structure is a lasting record of its evolutionary history, changing far more slowly than the gene sequence that encodes it. Using structures predicted by AlphaFold, we classified the domains in more than 120,000 proteins spanning the breadth of archaeal diversity and compared their folds with those of bacteria and eukaryotes. Archaea are not a hidden reservoir of novel structural designs: almost all of their folds are shared with the other domains of cellular life, and archaea occupy roughly 40% of known structural fold diversity. What distinguishes archaeal and eukaryotic cells is therefore not the invention of new folds, but new ways of combining and elaborating on an ancient shared repertoire.
Citation: Schaeffer RD, Pei J, Guo R, Zhang J, Medvedev K, Cong Q, et al. (2026) Domain classification of archaeal proteomes reveals conserved fold repertoire. PLoS Comput Biol 22(9): e1014188. https://doi.org/10.1371/journal.pcbi.1014188
Editor: Joanna Slusky, University of Kansas, UNITED STATES OF AMERICA
Received: April 1, 2026; Accepted: September 3, 2026; Published: September 18, 2026
Copyright: © 2026 Schaeffer et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All plots, associated data, and structure prediction outputs are in a Zenodo repository (https://doi.org/10.5281/zenodo.21656443).
Funding: This work was supported by grants from the National Institute of General Medical Sciences (GM147367 to R.D.S. and GM160468 to Q.C.; https://www.nigms.nih.gov/home), the National Institute of Allergy and Infectious Diseases (1K99AI180984-01A1 to J.Z.; https://www.niaid.nih.gov/), and the Welch Foundation (I-1505 to N.V.G.; I-2270-20260402 to Q.C.; https://welch1.org/). Computational resources were provided by NSF ACCESS (allocations MED230034 to J.Z. and MED240004 to N.V.G.; https://access-ci.org/) and TACC Lonestar6 (allocations MCB24018 to Q.C., MCB23014 to J.Z., and MCB26005 to R.D.S.; https://tacc.utexas.edu/). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Classifying the universe of protein domains requires structural data drawn broadly from across the tree of life. The Evolutionary Classification of protein Domains (ECOD) organizes domains into a hierarchy in which homologous superfamilies (H-groups) contain domains with common ancestry, while X-groups gather H-groups whose evolutionary relationship to each other has some evidence but remains uncertain [1, 2]. Below the homologous group, topology groups (or T-groups) divide homology groups among topologically distinct members who nevertheless retain homology; most H-groups contain only a single T-group. ECOD currently classifies over 2.7M protein domains from both experimental and predicted structures [3], but the taxonomic distribution of its source data is heavily skewed towards bacteria and eukaryotes. When structural data are drawn largely from a subset of organisms, relationships that exist in nature may go undetected, not because homology is absent, but because the intermediates needed to reveal it have not been sampled.
Archaea represent one of the most extreme cases of this sampling asymmetry. Though recognized as a primary domain of cellular life, whether as the third domain in the classical Woese framework or as the clade from which eukaryotes emerged in two domain models [4, 5], archaea account for ~2.5% of experimentally determined structures in the Protein Data Bank (PDB) and 2.7% of sequences in the AlphaFold Protein Structure Database (AFDB) [6]. This underrepresentation is not a minor gap: archaea span enormous phylogenetic and ecological diversity, from the reduced-genome DPANN lineages to the eukaryote-related Asgard archaea [7], and occupy environments ranging from hydrothermal vents and hypersaline lakes to temperate soils and ocean waters [8– 13] and the human gut microbiome [14]. Their proteins have evolved under selective pressures that differ from, and in many cases are more varied than, those shaping bacterial and eukaryotic proteomes, potentially driving exploration of regions of sequence and structure space that other lineages have not sampled. Prior phylogenomic analyses of domain structures across proteomes have noted archaea’s minimal structural repertoire [15], but these studies were limited to experimentally determined structures from cultivated organisms. Recent analysis of nearly 3,000 archaeal genomes found that even standard bioinformatic pipelines leave ~42% of archaeal genes uncharacterized, with long stretches of genes lacking even Pfam domain annotation [16], underscoring the gap between archaeal diversity and current classification tools. If novel protein folds remain to be discovered, archaeal proteomes are among the most likely places to find them.
Advances in protein structure prediction now make systematic surveys of underrepresented lineages feasible. AlphaFold2 and AlphaFold3 produce predicted structures with backbone accuracy approaching experimental methods for most globular proteins, though with reduced reliability for side chains and regions lacking homologous templates [17, 18]. Automated domain classification pipelines such as DPAM (Domain Parser for AlphaFold Models) can assign predicted structures to ECOD domains through iterative sequence and structural comparison [19]. These tools have been applied to proteome-scale datasets including 48 model organism and pathogen proteomes from the AlphaFold Database [20, 21], but no prior study has applied domain-level structural classification to archaeal proteomes at comparable scale. Meanwhile, structure-based comparison methods such as Foldseek detect remote homology well beyond the reach of sequence comparison [22], enabling both the recovery of relationships invisible to sequence methods and the identification of proteins with no detectable similarity to any known fold.
Here we present a systematic structural survey of 124,075 proteins from 65 archaeal classes spanning 21 phyla and all major archaeal lineages. Using existing structures from the AFDB and de novo AlphaFold3 predictions, we classified 204,758 protein domains against ECOD and annotated them against Pfam, producing a comprehensive archaeal domain classification. Structural clustering by Foldseek finds significant similarity for 63% of domains that are singletons by sequence comparison. Most significantly, we find that existing domain templates classify ~80% of archaeal proteins with high-confidence. Furthermore, systematic characterization of the unclassified fraction reveals that it is dominated by classification sensitivity limits and structure prediction quality rather than by novel structures. Because our sampling spans all major archaeal lineages rather than a few cultivated models, it also offers a pan-archaeal background against which to ask which structural components archaea share with eukaryotes, and whether folds proposed as Asgard-eukaryote connections are lineage-restricted or broadly archaea, a question we return to in considering the origins of eukaryotic cellular complexity. These results indicate that the protein fold repertoire at the single-domain level is broadly conserved across cellular life, and point toward family-level expansion, classifier sensitivity, and domain architecture diversity as the productive frontiers for structural classification.
Results
Dataset composition and structure prediction
The archaeal proteome dataset comprises 124,075 proteins spanning 65 Genome Taxonomy Database (GTDB) classes, 21 phyla, and 6 archaeal major groups; the full analysis workflow is summarized in Fig 1A. The dataset was assembled from three sources that differ in provenance: 71,866 proteins with existing AFDB structures, 22,883 with UniParc sequences but no AFDB structure (predicted here with AF3), and 29,326 predicted from Prodigal gene calls on metagenomic assemblies (sequence and structure both derived here, with AF3 structures, Fig 1B). We refer to the 71,866 AFDB proteins as the non-expanded set and the 52,209 AF3-predicted proteins (UniParc and Prodigal) as the expanded set, the structural coverage this study adds beyond AFDB. The six groups ranged from 6,543 (Deep-branching) to 32,066 proteins (Euryarchaeota-related), with DPANN (7,887), Thermoplasmatales (22,422), the TACK superphylum (23,641), and Asgard archaea (31,516) in between. The relative contribution of each data source varied by lineage: AFDB provided the majority of Euryarchaeota-related (83.4%) and TACK (78.9%) proteins, whereas Asgard archaea and Thermoplasmatales required mostly new structure predictions, with AFDB contributing only 35.1% and 26.3% of proteins, respectively.
(A) Analysis workflow. Target proteomes were selected from 65 Genome Taxonomy Database (GTDB) r220 classes and drawn from three sources of differing provenance: pre-existing structures from the AlphaFold Protein Structural Database (AFDB), UniParc sequences, and Prodigal gene predictions from metagenome-assembled genomes. Structures for the UniParc and Prodigal proteins were predicted de novo with AlphaFold3, giving 124,075 total predicted structures. Domains were classified using DPAM against ECOD v292, annotated against Pfam v38.2, and clustered by sequence (MMseqs2) and structure (Foldseek); this yielded 204,758 classified domains and, after successive filtering of the unclassified fraction, 2,852 dark-matter candidate proteins. (B) Dataset composition by major archaeal group and structure provenance. Stacked horizontal bars give the number of proteins whose structures were pre-existing in the AFDB (blue) versus newly predicted (red). Group sizes range from 6,543 (Deep-branching) to 32,066 (Euryarchaeota-related). The relative provenance contribution varies by lineage: AFDB supplies the majority of Euryarchaeota-related (83%) and TACK superphylum (79%) proteins, whereas Asgard archaea and Thermoplasmatales required mostly new predictions (AFDB contributing 35% and 26%, respectively). (C) Protein-length distributions. Violin plots with embedded box plots (median and IQR) for Swiss-Prot and the two archaeal provenance subsets: pre-existing and newly predicted in this study. Archaeal proteins are shorter than the Swiss-Prot reference, consistent with the compact genomes characteristic of prokaryotes; the two provenance subsets have near-identical length distributions indicating that the structure provenance introduces no systematic size bias. Proteins longer than 2,500 aa are excluded for display.
Archaeal proteins are shorter than those in Swiss-Prot, with a median length of 242 amino acids compared to 295 for Swiss-Prot (n = 542,378) (Fig 1C). The pre-existing (AFDB) and newly predicted subsets showed similar length distributions (median 240 vs. 245 amino acids), indicating that structure provenance did not introduce a systematic size bias. The pre-existing (AlphaFold2) and newly predicted (AlphaFold3) structures achieved comparable prediction confidence: the pre-existing set had a median per-residue pLDDT (predicted local-distance difference test, AlphaFold’s per-residue confidence, 0–100) of 89.3 compared to 86.2 for the newly predicted set, with similar residue coverage and domain density (S1 Fig). The modest difference in DPAM judge distributions (a categorical summary of DPAM confidence: good domain, simple topology, low confidence, or partial domain; see Methods), 79.9% high-confidence for the pre-existing set versus 72.4% for the newly predicted set, concentrated in the low-confidence category, does not compromise downstream analyses, which apply uniform quality filters. Of the 124,075 proteins, 116,275 (93.7%) have structure quality metrics, with an overall median pLDDT of 85.7.
ECOD classification landscape
DPAM assigned 204,758 domains across 115,623 proteins (S1 Table), when clustered by MMseqs2 or Foldseek, they were reduced to 123,169 domain sequence clusters and 52,341 domain structure clusters, respectively. Of these, 157,287 (76.8%) received high-confidence classifications (“well assigned”), representing strong matches to ECOD domain templates. The most abundant structural families are broadly distributed across archaeal lineages (Fig 2A). The top 15 H-groups, including HTH domains (H-group 101.1), P-loop NTPases (H-group 15.1), and Rossmann-related folds (H-group 2003.1), account for 62,439 of 157,287 high-confidence domain assignments (39.7%), with representatives present in all six major archaeal groups. Archaeal domains span a substantial fraction of known structural diversity (Fig 2B): 987 of 2,457 ECOD X-groups (40.2%), 1,583 of 3,716 H-groups (42.6%), and 1,751 of 3,953 T-groups (44.3%). For comparison, Swiss-Prot (representing all three domains of life) covers 1,874 X-groups, 2,851 H-groups, and 3,060 T-groups. Single-domain proteins predominate (63,473 of 115,623; 54.9%), followed by two-domain proteins (32,370; 28.0%), with progressively fewer multi-domain architectures up to a maximum of 20 domains per protein (Fig 2C).
(A) Top 15 ECOD H-groups (homologous superfamilies) by domain count, restricted to well-assigned domains. Stacked bars show contributions from each major archaeal lineage. HTH domains, P-loop NTPases, and Rossmann-like folds are the most abundant structural families, with broad distribution across all archaeal groups. The top 15 H-groups account for 62,439 of 157,287 well-assigned domains (39.7%). (B) ECOD hierarchy coverage comparison. Archaeal domains span 987 X-groups (40% of 2,457 in ECOD), 1,583 H-groups (43% of 3,716), and 1,751 T-groups (44% of 3,953). Swiss-Prot covers 1,874 X-groups, 2,851 H-groups, and 3,060 T-groups. A single domain of life thus captures over 40% of known structural diversity. (C) Distribution of domains per protein for all 115,623 proteins with domain assignments (median = 1). Single-domain proteins predominate (54.9%), with 28.0% having two domains and progressively fewer multi-domain architectures.
To assess the global impact of expanded archaeal sampling on ECOD taxonomic occupancy, we compared the distribution of H-groups across the three superkingdoms before and after incorporation of the archaeal proteome dataset (Fig 3). Superkingdom occupancy was taken from ECOD (v292) annotations, the classified domains assigned to each H-group in bacteria, eukaryota, and archaea. The most prominent change is a marked expansion of the universal set: H-groups present in all superkingdoms increased from 746 to 1,326. In parallel, the number of H-groups assigned only to eukaryota and bacteria decreased from 1,329–879, while eukaryote-only and bacteria-only categories also contracted. By contrast, archaea-only H-groups increased only marginally, from 34 to 39. These shifts indicate that many H-groups previously interpreted as absent from archaea were in fact unsampled, and that archaeal domains occupy a far broader fraction of the shared cellular fold repertoire than suggested by earlier datasets.
(A,B) Venn diagrams showing the distribution of ECOD H-groups across Eukaryota (E), Bacteria (B), and Archaea (A) in the full ECOD database before (A) and after (B) incorporation of the archaeal proteome dataset. (C) Bar plot showing the change in H-group counts for each superkingdom occupancy category after archaeal expansion.
To place this shift in context, we first asked how many archaeal folds are visible in a single model organism. Considering only Methanocaldococcus jannaschii, historically the only archaeon with a fully classified proteome, recovers 542 H-groups shared with other domains and only 13 archaea-specific H-groups. Extending to all archaea currently in ECOD raises these to 746 and 34, respectively. Our full archaeal classification raises them further to 1,326 and 39. The apparent scarcity of archaea-specific folds is therefore largely a sampling effect: broader sampling reassigns folds from lineage-restricted to shared categories rather than uncovering novel folds.
Even with broad archaeal sampling, the number of archaea-exclusive H-groups remains small. The few additional archaeal-specific folds identified after expansion are mostly associated with niche or lineage-restricted functions, including CRISPR defense systems, mobile-element or viral-related proteins, and uncharacterized families, rather than core metabolic or informational processes. At the same time, many H-groups previously assigned to eukaryotes, bacteria, or their intersection are reassigned to shared categories, especially the universal set.
These observations show that expanded archaeal sampling mainly reveals previously unsampled members of existing fold families rather than introducing substantial new fold diversity. Archaeal proteomes therefore appear to be constructed largely from a common cellular fold repertoire, with lineage-specific innovation arising mainly through specialization and diversification within a limited set of structural frameworks.
Domain repertoire across archaeal lineages
The six major archaeal groups share a common domain repertoire that varies in size but not kind. Mapping high-confidence H-group content onto the accepted phylogeny of the six groups (S2 Fig) shows a universal core of 551 of 1,583 H-groups (35%) present in all six, which constitutes the majority of each group’s repertoire, so repertoire size tracks genome scope through shared accessory folds rather than the emergence of new fold classes. Among well-assigned (good-domain) H-groups, counts vary from 766 (DPANN) to 1,227 (Euryarchaeota-related), largely tracking genome size; 97% of DPANN H-groups are shared with other archaea. The 23 DPANN-unique H-groups are mostly singletons (21 of 23) drawn from families whose ECOD representation is otherwise bacterial or eukaryotic (e.g., chitin-binding domains, Kazal-type protease inhibitors), consistent with either horizontal gene transfer or classification artifacts in highly-reduced genomes, though individual cases have not been systematically verified. Turning to folds shared specifically between archaea and eukaryotes, 143 H-groups are occupied by both but absent from bacteria; 118 are supported by a confidently assigned eukaryotic domain, while the remaining 25 rest on low-confidence eukaryotic assignments, simple or partial topology matches DPAM did not confidently resolve, and are excluded from the homology-supported count. This set includes the recognizable informational core shared between the two domains of life: the archaeal- and eukaryotic-specific RNA polymerase subunits Rpb5 and Rpb10, tRNA-splicing endonucleases, DNA replication initiators (Cdc45, cdc21/45, GINS), and translation factors (eIF2alpha and elongation factors), now resolved across all major archaeal lineages rather than through a single model organism [23].
Pan-archaeal scope also recontextualizes recent reports of eukaryotic signature proteins (ESPs) in Asgard archaea. Our classification recovers known ESP families in Asgard lineages, consistent with recent structural analyses [7, 24] and characterizations of defense system repertoires shared between Asgard archaea and eukaryotes [25]. However, several protein families reported as Asgard-eukaryote connections, including the major vault protein (MVP; H-group 3529.1), classified in ECOD since 2014, are present with high-confidence across all six major archaeal groups, with Asgard accounting for only 16% of MVP domain assignments. The distinction is that folds proposed elsewhere as Asgard-eukaryote signature are not uniformly Asgard-specific: several, most clearly MVP, are distributed across all major archaeal groups. Whether a given signature fold is truly Asgard-specific or broadly archaeal can therefore only be established against the full archaeal repertoire. This systematic accounting is beyond the present scope but is a target for follow-up.
Annotation landscape and the unclassified fraction of domains
We assessed the completeness and reliability of ECOD classifications by comparison with Pfam annotations. Pfam concordance varied systematically by DPAM judge category (Fig 4A): high-confidence assignments showed a 73.8% Pfam mapping rate (116,115 of 157,287), declining through partial domain (36.7%), low confidence (25.1%), and simple topology (12.1%). The declining Pfam rate across judge categories is consistent with DPAM's confidence rankings reflecting genuine differences in sequence divergence from characterized families, not arbitrary thresholds.
(A) Pfam mapping rate by DPAM classification judge. Well-assigned domains show the highest Pfam concordance (73.8%, 116,115 of 157,287), followed by partial (36.7%), low confidence (25.1%), and simple topology (12.1%). The declining Pfam rate across judge categories is consistent with increasing sequence divergence from characterized families. (B) Domain annotation landscape comparing archaeal and Swiss-Prot proteomes. Four categories are shown: ECOD + Pfam (both structural and sequence classification), ECOD only (structural match without Pfam), Pfam only (sequence match without high-confidence ECOD), and Orphan (neither). Archaea show 56.7% ECOD + Pfam, 20.1% ECOD only, 4.9% Pfam only, and 18.2% orphan domains. Swiss-Prot shows 75.2% ECOD + Pfam and only 7.7% orphan domains. The higher orphan fraction in archaea reflects their underrepresentation in current classification databases. (C) Novel domain cluster size distribution (Tier 2 orphan domains: no Pfam hit, sub-threshold DPAM assignment). Foldseek structural clustering of 37,338 orphan domains yields 29,817 clusters, of which 2,097 are multi-member. The CCDF shows the cumulative number of domains in clusters at or above each size, with the largest clusters containing hundreds of structurally similar orphan domains, candidates for novel or highly divergent domain families.
A four-way annotation landscape (Fig 4B) reveals that 56.7% of archaeal domains carry both a high-confidence ECOD classification and a Pfam hit, 20.1% carry ECOD only, 4.9% carry Pfam only, and 18.2% lack both (“orphan” domains). In our classification of Swiss-Prot proteins [26], orphans constitute only 7.7% of domains. The elevated orphan fraction in archaea, more than double the Swiss-Prot rate, raises a question that motivates the remainder of this analysis: does this gap represent genuine structural novelty among archaeal proteins, or is it a consequence of classifying deeply divergent sequences against a reference database built largely from bacterial and eukaryotic structures?
To begin addressing this question, we examined the 25,965 proteins that lack any high-confidence domain assignment (Fig 4C). Of these, 67.5% have at least one DPAM domain hit below the high-confidence threshold; they are sub-threshold matches to known folds, not structurally uncharacterized. The remaining 32.5% lack any DPAM domain assignment whatsoever. These domain-free proteins are characterized in detail below; as we show, the majority are explained by low-confidence structure predictions and short protein length rather than by novel folds.
Structural clustering reveals relationships beyond sequence detection
To explore relationships beyond the reach of sequence-based classification, we clustered at four levels: protein and domain sequence clusters (PSC, DSC) using MMseqs2, and protein and domain structural clusters (PXC, DXC) using Foldseek. Sequence-based clustering identified a high proportion of singletons: 53.2% at the protein level and 51.7% at the domain level (Fig 5A). Structure-based clustering sharply reduced these fractions to 15.2% and 20.2%, respectively. At the domain level, 66,399 sequence singletons were rescued into multi-member structural clusters, a 63% recovery rate. The greater sensitivity of structural clustering is evident in cluster size distributions (Fig 5B): structural clusters consistently show heavier tails than their sequence counterparts on log-log plots. Many DXC clusters unify members from multiple DSC clusters (Fig 5C), with 9,800 DXC clusters bridging two or more DSC clusters, indicating that archaeal protein families harbor far more internal diversity than sequence comparison alone can capture. Clustering parameters (E-value 0.001, 50% coverage, TM-align mode) were chosen for sensitivity in detecting remote relationships; structural similarity at these thresholds indicates shared topology but does not by itself establish common evolutionary origin, especially for abundant folds such as Rossmann-like and TIM barrel domains where convergence is well documented.
(A) Singleton fractions for sequence-based versus structure-based clustering at both protein and domain levels. Structure-based clustering recovers 63% of domain-level sequence singletons into multi-member clusters. (B) Complementary cumulative distribution functions (CCDFs) of cluster sizes across all four clustering dimensions on a log-log scale. Structural clusters show heavier tails than sequence clusters. (C) Number of distinct sequence clusters unified within each structural cluster. At the domain level, 9,800 structural clusters contain members from two or more sequence clusters.
Characterizing the structural dark matter
We focused on the 8,452 proteins lacking any DPAM domain assignment and applied successive filters to identify genuinely novel structures (Fig 6A). The largest filter is structure quality: 74.0% have a mean pLDDT below 70, indicating disordered regions or poor structure predictions unlikely to harbor classifiable domains. The dramatic separation in pLDDT distributions between classified proteins (median 89.6), sub-threshold proteins (median 76.7), and domain-free proteins (median 62.6) confirms that unclassifiable proteins are chiefly those for which AlphaFold did not produce confident structures (Fig 6B).
(A) Dark matter filtering pipeline. Of 8,452 proteins lacking any DPAM domain assignment, successive filters remove those with low-confidence structures (pLDDT < 70; 74%), lengths below template sensitivity (< 100 aa; 14%), structural singletons (1%), membership in clusters containing classified members (7%), and membership in clusters with sub-threshold domain hits (3%). The residual, 36 proteins in 20 fully dark clusters, represents the irreducible core of potentially novel folds, constituting 0.03% of the 124,075-protein dataset. Of these 20 clusters, 18 are pairs (size 2); only PXC_023126 (10 members, 8 lineages; maps to Pfam DUF2769, present in ECOD but missed by DPAM) and PXC_008059 (10 members, 10 lineages spanning both Asgardarchaeota and Thermoproteota; 80% AFDB-source; no Pfam or InterPro annotation) are substantial. PXC_008059 represents the strongest candidate for a genuinely novel fold. (B) pLDDT distributions by classification category. Proteins with high-confidence ECOD classification (blue; median 89.6, n = 94,229) are strongly enriched for well-folded structures, while proteins lacking any domain assignment (gray; median 62.6, n = 5,854) are mostly disordered or poorly predicted. Sub-threshold proteins (orange; median 76.7, n = 16,148) occupy an intermediate range. Domain rescue by structural clustering: of 10,148 sub-threshold domains in structurally consistent clusters, 82% agree with the X-group consensus of their classified cluster neighbors, rescuing 3,463 proteins from unclassified to classifiable status.
The unclassified fraction distributes unevenly across lineages. Classification rates range from 82.5% (Euryarchaeota-related) to 74.7% (DPANN). DPANN has the highest no-domain rate (9.0%) but the lowest median pLDDT among domain-free proteins (53.9), confirming that its unclassified fraction reflects structure prediction quality in small, fragmentary genomes rather than novel folds. DPANN contributes only 6 of the final well-folded dark candidates after filtering. Asgard contributes the most (45% of candidates), and Prodigal-source proteins are overrepresented relative to their dataset share, consistent with gene prediction artifacts inflating apparent novelty. To confirm that this residual reflects intrinsic disorder and prediction quality rather than concealed structural novelty, we predicted disorder directly from sequence with metapredict V3 [27] for all 8,760 domain-free proteins and a 2,500-protein high-confidence classified control (S3 Fig). The control was ordered (94.2% with <20% predicted-disordered residues), whereas the domain-free fraction was bimodal, with a distinct disordered population (≥ 50% of residues; 24.3%) essentially absent from classified proteins. Sequence-predicted disorder correlated only weakly with structure confidence (Pearson r = -0.19), indicating that disorder and poor prediction are not causally linked in this protein set. The proteins carried forward as novel-fold candidates were almost entirely sequence ordered (288 of 295; median 0% disorder), confirming that the curation pool comprises genuine domains rather than disorder artifacts.
Of the 2,196 well-folded domain-free proteins, 1,147 are shorter than 100 amino acids – below the effective sensitivity of template-based domain classification – and 125 are singletons in structural clustering with no independent structural support. This leaves 924 well-folded, reasonably sized proteins in non-singleton Foldseek clusters. Most of these cluster with proteins that do carry ECOD domains: 633 are in clusters containing at least one high-confidence domain member (i.e., they are classifiable by transitivity), and 255 are in clusters with sub-threshold domain members (i.e., also have partial classification signal). The residual consists of a small number of proteins distributed across fully dark clusters, clusters where no member has any ECOD domain assignment at any confidence level. These clusters, which include a mix of data provenance and phylogenetic breadths, are candidates for genuine structural novelty and require further curation to distinguish bona fide novel folds from gene prediction artifacts and extreme classification sensitivity failures.
We also assessed the reliability of DPAM classifications using structural clustering as a validation (Fig 5C). Of 3,662 DXC clusters containing both high-confidence and sub-threshold domain members, 91.4% show consistent X-group consensus among their high-confidence members. Among the 10,148 sub-threshold domains in these consistent clusters, 81.5% carry a prior X-group assignment that agrees with the cluster consensus, 16.7% disagree, and 1.7% receive a new X-group assignment that they previously lacked. The disagreements are concentrated at multi-domain boundaries involving HTH motifs, where domain parsing ambiguity rather than incorrect fold assignment is the likely cause. We restricted rescue to clusters with unanimous high-confidence consensus, a conservative design reflecting the fact that structural cluster membership alone is not sufficient evidence for homology assignment, especially for common folds where structurally similar domains may belong to different evolutionary lineages. Under this criterion, 3,463 proteins are reclassifiable, bringing the effective classification rate from ~79.1% to ~81.9%.
Manual analysis of dark archaeal protein domains revealed several notable cases. These domains were classified by their overall secondary-structure content, three-dimensional architecture, and size. Among them, seven small cysteine-rich domains are likely to bind metals such as zinc (Fig 7A). Two of these (EPP00014805 and EPP00002661) contain eight conserved cysteines each and may therefore coordinate multiple metal ions (Fig 7A, leftmost two structures). Together with another small domain (third structure from the left in Fig 7A, ECOD: EPP00018990), they may represent previously unrecognized metal-binding folds. In contrast, four other small cysteine-rich domains show structural similarity to known zinc fingers. One domain (EPP00027814) appears to be a remote homolog of the zf-HC3 family in Pfam (PF16827), whereas the other three resemble treble-clef zinc fingers [28], which are defined by a zinc knuckle (a β-hairpin containing two zinc-binding residues) followed by an alpha-helix whose N-terminus contributes two additional zinc-binding residues. This treble-clef core is elaborated by distinct surrounding secondary-structure elements, including, in one case (EPP00000606), a C-terminal transmembrane helix (rightmost structure in Fig 7A).
(A) Cysteine-rich metal-binding domains. Conserved cysteine and histidine residues that likely coordinate metal ions are shown in pink. (B) Three ɑ + β two-layered domains with potential novel folds. (C) Two ɑ-helical domains likely involved in defense. (D) A domain with complex topology representing a potential novel fold. (E) Three β-barrel domains. (F) Two ɑ + β two-layered domains. The right structure adopts a Rossmann-like fold (ECOD X-group 2003). The left structure is modeled as a homodimer with two zinc ions using AlphaFold3; one subunit is colored in a rainbow scheme and the other in gray. Side chains of conserved metal-binding and catalytic residues are shown as sticks. (G) A β-sandwich domain with a potential novel fold.
Other architectures include ɑ + b two-layer folds (a single b-sheet with a-helices on one side; 6 cases), ɑ-helical arrays or bundles (5 cases), β-barrels (2 cases), ɑ + β three-layer folds (2 cases), β-sandwiches (2 cases), and ɑ + β complex topology (2 cases). Several domains appear to represent novel folds, including three ɑ + β two-layered domains (Fig 7B, A0A060HJU8, EPP00026547 and EPP00002661), one ɑ + β domain with complex topology (Fig 7D, EPP00003247), two β-barrels (Fig 7E, left and middle structures, accession: EPP00010451 and EPP00001136), and one β-sandwich domain (Fig 7G, EPP00002877).
Detailed manual analysis of sensitive sequence and structural similarity searches further suggests that some of these domains may be remote homologs of known protein families. For example, one ɑ-helical domain (accession: EPP00023006) was identified as a distant homolog of the C-terminal region of DndB (DndB_C; Fig 7C, left structure, HHpred score: 98.51), which is involved in DNA phosphorothioate modification [29]. Another alpha-helical bundle domain (EPP00023009) shows similarity to the HEPN (ECOD H-group 601.7) superfamily of endoribonucleases [30] and contains several conserved residues that are likely to form the active site (Fig 7D). In addition, one case (UPI00355E9374) represents a case of duplication of β-clip fold (ECOD X-group 70) domains, which are found in prokaryotic phage capsid decoration (cement) proteins located on the outside surface of the icosahedral capsid [31].
We also identified a domain with a Rossmann-like fold (ECOD X-group 2003, Fig 7F, left structure, UPI003165088D) that does not exhibit detectable sequence similarity to any known domains. Structural searches using DaliLite identified domains in the HAD superfamily of phospho-hydrolases [32] as the closest matches, characterized by short β-strands and two ɑ-helices in the Rossmann-fold crossover region. However, this domain lacks the canonical metal-binding active-site residues of HAD enzymes and contains additional C-terminal secondary-structure elements, including an antiparallel β-strand. Thus, its potential evolutionary relationship to HAD domains remains unclear.
Another ɑ + β three-layered protein domain (Fig 7F, right structure, UPI00355F993C) shows structural similarity to zincin domains and contains the characteristic HEXXH motif of zincin-like metalloproteases [33]. However, in monomeric models, the third conserved acidic residue located in the final ɑ-helix is positioned distal to the HEXXH motif, unlike in canonical zincins. When modeled as a homodimer (one subunit colored rainbow and the other colored gray in Fig 7F right structure) using AlphaFold3 in the presence of two zinc ions, the acidic residue from the adjacent monomer is positioned close to the HEXXH motif, completing the metal-binding site at the dimer interface. To our knowledge, this represents a unique arrangement among zincin-like metalloproteases (ECOD X-group 2498).
The archaeal fold repertoire in context
Combining direct classification with structural clustering rescue, the accounting of the 124,075 archaeal proteins is as follows. A total of 98,110 proteins (79.1%) are classified directly by DPAM with high confidence, and an additional 3,463 (2.8%) are reclassifiable through structural cluster transitivity. The sub-threshold fraction, 17,513 (14.1%) proteins with domain matches below the confidence threshold, represent known homologous groups detected at reduced sensitivity, not structural novelty. Among the domain-free proteins, the majority are explained by low-confidence structure proteins (5.0%), short length below classification sensitivity (0.9%), or lack of independent structural support (0.1%). A small number of candidates for genuine structural novelty remain under investigation.
Existing homologous groups thus account for virtually all confidently predicted archaeal protein structures. The sub-threshold fraction, the single largest component of the unclassified proteome, represents a classification sensitivity gap rather than a structural novelty gap: these proteins have detectable similarity to known folds, just below the threshold for high-confidence (and automated) assignment.
MCR and MVP illustrate lineage-specific and pan-archaeal domain distributions
To illustrate the biological information captured by pan-archaeal domain classification, we highlight two protein families with contrasting phylogenetic distributions (Fig 8). Methyl-coenzyme M reductase (MCR) exemplifies a lineage-restricted, archaea-exclusive enzyme family. The two MCR domains (H-groups 304.35 and 154.1) exhibit zero bacterial occupancy and, at the homology-supported F-group level, zero eukaryotic members in the ECOD database; two eukaryotic domains assigned through topology-based promotion to these H-groups represent misclassifications currently flagged for manual curation.. Within our dataset, MCR domains are present in classical methanogens (Methanosarcina, Methanococcus, Methanopyrus), but also in organisms that use MCR-like enzymes for anaerobic alkane oxidation rather than methanogenesis: including Ca. Syntrophoarchaeum butanivorans (anaerobic butane oxidizer; 20 MCR domains, the most in our dataset) and Ca. Hadarchaeales (deep-branching archaea with proposed hydrocarbon metabolism). Bathyarchaeota carry MCR homologs whose function remains debated. The conserved two-domain architecture, an N-terminal ɑ/β domain (304.35) packed against a C-terminal helical domain (154.1), is maintained across these functionally diverse lineages, with size variation (262–607 aa) but not domain organization (Fig 8B). MCR domains are absent from Asgard, DPANN, and non-methanogenic Euryarchaeota, consistent with a metabolism-linked rather than phylogeny-linked distribution. Domain classification identifies the shared structural scaffold but cannot distinguish the reaction direction; the same two-domain architecture supports both methane production and alkane activation.
(A) Major vault protein (MVP) representatives from five (of six) major archaeal groups, colored by ECOD domain: green = H-group 3529.1 (MVP repeat), blue = H-group 230.5 (SPFH/Band 7), orange = lineage-variable C-terminal domains, gray = unassigned. The conserved repeat-shoulder-tail architecture spans Asgard (A0A9Y1BM57, 392 aa), DPANN (UPI001A0E75C9, 322 aa), Euryarchaeota (D4GYN0, 396 aa), TACK (B1L689, 328 aa), and deep-branching archaea (UPI001765BDB6, 317 aa). (B) Methyl-coenzyme M reductase (MCR) representatives from functionally diverse lineages, colored by ECOD domain: blue = H-group 304.35 (MCR N-terminal), red = H-group 154.1 (MCR C-terminal), gray = unassigned. The conserved two-domain architecture is maintained across a Syntrophoarchaeum alkane oxidizer (A0A1F2P4I2, 460 aa), deep-branching Hadarchaeales (UPI003168A123, 607 aa), and a TACK methylotroph (L0L0H0, 571 aa), with a single-domain variant from Methanoliparum (A0A520KRH7, 262 aa). MCR domains are archaea-exclusive (zero bacterial, zero eukaryotic in ECOD) and distributed by metabolism rather than phylogeny. (C) Eukaryotic MVP from rat liver vault (PDB 2ZUO), shown as a single chain (845 aa) and as the 13-chain half-vault assembly, with the same color scheme. The eukaryotic monomer elaborates the archaeal architecture with 8 tandem MVP repeats (green), a retained SPFH shoulder (blue), and a helical cap domain (orange, H-group 3530.1).
The major vault protein (MVP) illustrates the opposite pattern. MVP was recently reported as an Asgard-eukaryote structural connection by Köstlbacher et al. [24], who identified Asgard proteins with reciprocal best structural hits to eukaryotic MVP. However, the MVP structural repeat (H-group 3529.1) has been classified in ECOD since 2014 (PDB 2ZUO, rat liver vault). Our pan-archaeal classification identifies 32 high-confidence MVP domains across all six major archaeal groups, with Asgard accounting for only 5 (16%). The conserved multi-domain architecture; an N-terminal MVP repeat (3529.1), a central SPFH/Band 7 domain (230.5), and lineage-variable C-terminal domains, is present from DPANN to Asgard (Fig 8A). In the eukaryotic vault (2ZUO), the same architecture is elaborated: 8 tandem MVP repeats replace the single archaeal repeat, the SPFH shoulder is retained, and a long helical cap domain (3530.1) extends into the vault particle interior (Fig 8C). The archaeal proteins thus represent a compact version of the eukaryotic vault monomer. Whether archaeal MVPs assemble into vault-like particles is unknown, but the domain-level conservation suggests that the vault architecture predates the archaeal-eukaryotic divergence rather than originating as an Asgard-specific innovation. The distinction between Asgard-specific and pan-archaeal distribution is only visible with the breadth of sampling that pan-archaeal classification provides.
Discussion
The central finding of this study is that the protein fold repertoire at the single-domain level is broadly conserved across cellular life. Archaea, the most phylogenetically distant and structurally undersampled domain, are built from the same set of domain-level building blocks as bacteria and eukaryotes. Of 124,075 proteins spanning all major archaeal lineages, ~ 80% receive high-confidence domain classifications directly, and systematic characterization of the unclassified fraction reveals that it is dominated by classification sensitivity limits and structure prediction quality rather than by novel structures.
The unclassified fraction represents sensitivity, not novelty
The elevated orphan fraction in archaeal domains (18.2% versus 7.7% in Swiss-Prot) initially suggests extensive uncharacterized structural diversity. However, our analysis reveals that this gap is driven mainly by classification sensitivity. The largest contributor is structure prediction quality: 74% of proteins lacking any domain assignment have mean pLDDT scores below 70, indicating disordered regions or unreliable predictions rather than well-folded domains of unknown type. Among well-folded, reasonably sized domain-free proteins, structural clustering connects most to proteins that do carry DPAM classifications, either directly (classifiable by transitivity) or through sub-threshold matches (partial signal). The funnel from 8,452 domain-free proteins to a small number of uncharacterized candidates reflects a series of straightforward explanations, poor predictions, short sequences, known folds at sub-threshold confidence, that collectively account for the majority of the unclassified fraction.
The 14% sub-threshold fraction merits particular attention. These proteins have DPAM domain matches to known folds, just below the confidence threshold used for high-confidence classification. Structural clustering validates this interpretation: in mixed clusters containing both high-confidence and sub-threshold members, 81.5% of sub-threshold domains carry X-group assignments that agree with the cluster consensus. The dominant issue is calibration, the sensitivity of template-based classification for deeply divergent archaeal sequences, not the absence of matching templates. This finding suggests that the next gains in archaeal classification will come from improving classifier sensitivity rather than from discovering new folds.
Conservation of the topology repertoire across cellular life
The question of whether the protein fold universe is finite and mappable has recurred across multiple eras of structural biology. Structural genomics initiatives of the early 2000s consistently found that new folds appeared at a diminishing rate [34– 36], and estimates of the total number of domain-level folds converged on a range of 1,000–10,000 [37– 39]. The predicted structure revolution has revisited this question at a qualitatively different scale: where structural genomics solved thousands of structures per year, AlphaFold now provides models for hundreds of millions of proteins. At each stage of expansion, from PDB-only to predicted structures from 48 model organisms [3, 26] to the archaeal proteomes surveyed here, the proportion of genuinely novel folds discovered has been small relative to the expansion of coverage within known fold families.
Our results extend this pattern to the most phylogenetically extreme test available within cellular organisms. Archaea include lineages (Asgardarchaeota, DPANN, deep-branching thermophiles) with no close relatives among well-characterized organisms. If large regions of fold space were accessible only through deeply divergent lineages, archaeal proteomes should reveal them. That existing templates instead cover archaeal proteins at rates comparable to bacterial and eukaryotic proteomes suggests that the domain-level fold repertoire is shared across cellular life, a statement about the universality of the structural building blocks, not an enumeration of all possible folds.
Several important caveats apply. First, this conclusion is restricted to the single-domain level; the combinatorial space of domain arrangements is far larger than the space of individual folds and remains poorly characterized. Second, and critically, structure prediction methods are inherently conservative: AlphaFold is trained on known structures and produces confident predictions most readily for proteins that resemble its training data. The pLDDT filter that removes 5% of proteins as “disordered or poorly predicted” may also exclude proteins whose structures are genuinely novel but fall outside the prediction model's confident regime. We cannot fully distinguish between “no novel folds exist” and “our methods cannot confidently predict novel folds.” The strongest version of our claim is comparative: archaeal proteins are classified at the same rate as other cellular proteomes, indicating that archaea are not a reservoir of unexplored fold diversity, but this does not preclude the existence of novel folds among proteins for which no confident prediction is available. A systematic test of whether X-group boundaries, which represent uncertain evolutionary relationships between homologous superfamilies, are robust to archaeal diversity found no new cross-X-group relationships among archaeal domains, consistent with the overall conservation finding.
Implications for eukaryotic origins
Our results also bear on the relationship between archaea and eukaryotes, though the fold data alone do not adjudicate between the two- and three-domain models of the tree of life. At the single-domain level we recover a substantial archaeal-eukaryotic shared repertoire: 143 H-groups are shared between archaea and eukaryotes but absent from bacteria, and this set includes the recognizable information-processing core (archaeal- and eukaryotic-specific RNA polymerase subunits such as Rpb5 and Rpb10, tRNA-splicing endonucleases, and replication and translation factors) expected of a shared informational heritage, now resolved across all major archaeal lineages rather than simply M. jannaschii. We also recover known eukaryotic-signature protein families across Asgard lineages. However, we are cautious about reading Asgard-eukaryote fold sharing directly as phylogenetic signal: our pan-archaeal sampling shows that several folds proposed as Asgard-eukaryote connections, most clearly the major vault protein, are distributed across all major archaeal groups rather than restricted to Asgard. Therefore, we consider our classification to be consistent with an archaeal ancestry of eukaryotic cellular complexity, but not definitive to it. Domain classification records which structural building blocks are shared, whereas distinguishing vertical inheritance from within Asgard from deep shared ancestry is a problem of phylogenetic inference (marker-gene trees, taxon sampling, and model adequacy) that structural classification cannot settle. The tractable, structural classification question adjacent to this debate is whether the archaeal-eukaryotic shared repertoire is uniformly pan-Asgard or concentrated in particular Asgard clades. Resolving it requires Asgard sampling deeper than the one-representative-proteome-per-class design used here. The recently expanded Asgard genomic catalog (e.g., the 936-genome collection of Köstlbacher et al.) offers exactly that depth, and extending domain-level classification across it is the natural next step toward placing the archaeal-eukaryotic fold repertoire in an explicit evolutionary framework.
Data provenance and confidence
The phylogenetic breadth that makes this dataset a strong test of fold repertoire conservation is itself dependent on metagenome-assembled genomes. Of the 65 GTDB classes in our dataset, ~ 60% are represented primarily or exclusively by MAG-derived organisms, including the Asgard archaea, much of DPANN, and most deep-branching lineages. Without metagenomics, systematic structural coverage of archaeal diversity would be limited to the ~ 18 cultivated classes.
The three data sources carry different levels of confidence. Proteins from the AlphaFold Database (39 classes, 58% of the dataset) have curated gene models and structures validated through the AFDB pipeline. UniParc-source proteins (11 classes, 18%) have established sequences but require de novo structure prediction. Prodigal-source proteins (15 classes, 24%) carry compound uncertainty: gene boundaries, protein sequences, and structures are all derived computationally from metagenomic assemblies of varying quality. This provenance gradient is visible in the dark matter analysis, where Prodigal-source proteins are overrepresented among uncharacterized candidates relative to their share of the full dataset, consistent with gene prediction artifacts inflating the apparent novelty. The conservation finding is robust to data source: even restricting to AFDB-source proteins alone, the fraction of well-folded structural orphans remains below 0.1%.
Open frontiers
These results point to several productive directions for structural classification. The most immediate is family-level expansion. The ~ 1,000 fold-level building blocks may be broadly shared across cellular life, but family-level membership, the sequence and structural diversity within each fold, is where lineage-specific adaptation lives. Comprehensive classification of the full predicted structure universe, now encompassing hundreds of millions of proteins, is the natural next step for mapping this diversity.
Viral proteomes represent a qualitatively different frontier. Viruses evolve under fundamentally different constraints from cellular organisms, including de novo gene origination and capsid architecture, and are not bound by the same evolutionary continuity that links archaeal, bacterial, and eukaryotic folds. If genuinely novel folds exist beyond the cellular repertoire, viral proteomes are the most promising place to look.
Multi-domain architectures constitute another open question. Our analysis is restricted to the single-domain level, but the combinatorial space of domain arrangements is far larger than the space of individual folds. Domain co-occurrence rules, architecture evolution, and lineage-specific domain combinations remain poorly characterized and represent a substantial source of protein functional diversity.
Finally, classifier sensitivity is a tractable engineering problem. The 14% sub-threshold fraction represents the largest opportunity for improving archaeal classification coverage. The high-confidence threshold used for automated ECOD accession reflects an operational trade-off between classification accuracy and throughput, and many sub-threshold matches are likely correct but fall below the confidence level at which assignments can be made without manual review. Structural clustering offers one path to promoting these assignments, our cluster-based validation suggests that most of these domains are correctly assigned at the X-group level, but using structural similarity as a direct proxy for evolutionary classification carries its own risks, especially for superfolds where convergent evolution can place unrelated domains in the same structural cluster. The path forward likely involves combining expanded template libraries, improved profile-based search for divergent sequences, and cluster-informed classification with appropriate caution about the distinction between structural similarity and homology.
Conclusions
Systematic structural classification of archaeal proteomes shows that the protein fold repertoire at the single-domain level is broadly conserved across all major cellular lineages. The gap between archaeal and well-characterized proteomes reflects classification sensitivity for divergent sequences, not an abundance of novel structures. The priorities for structural classification accordingly shift from fold discovery toward family-level expansion, sensitivity improvements, and characterization of the combinatorial diversity of domain architectures that these conserved building blocks assemble into.
Methods
Genome selection and protein dataset assembly
Archaeal genomes were selected from the Genome Taxonomy Database (GTDB) release r220 [40], which contains 11,918 archaeal genome assemblies (1,178 cultivated and 10,740 metagenome-assembled genomes [MAGs]). To complement existing AlphaFold Database (AFDB) coverage, we excluded genomes already represented in AFDB (8,800 MAGs encompassing ~14.3 million proteins) and targeted the remaining genomes for download from NCBI. Protein sequences were obtained for 954 genomes: 367 cultivated genomes (123 from UniProt reference proteomes, 244 from NCBI GenBank/RefSeq) and 587 MAGs. MAGs were classified by CheckM quality metrics: 236 as high-quality (completeness ≥90%, contamination ≤5%) and 351 as medium-quality (completeness ≥50%, contamination ≤10%).
From this pool, a representative working set of 65 GTDB classes was curated to balance phylogenetic coverage across known archaeal lineages, inclusion of both cultivated organisms and MAGs, and representation of phylogenetically distinct groups that are undersampled in existing structural databases. These 65 classes span 21 phyla and were organized into 6 operational groups based on established archaeal phylogeny: Asgard, DPANN, TACK superphylum, Euryarchaeota-related, Deep-branching, and Thermoplasmatales. The working set comprises 124,075 proteins drawn from three sources that differ in provenance and confidence. The first source (39 classes, 71,866 proteins) consists of proteins with existing AlphaFold2 structures in the AFDB, representing a mix of cultivated reference organisms with curated gene models and MAG-derived organisms already accessioned by AFDB. The second source (11 classes, 22,883 proteins) consists of proteins with existing sequences in UniParc but no AFDB structures; for these, we predicted structures de novo using AlphaFold3. The third source (15 classes, 29,326 proteins) consists of proteins from lineages with neither AFDB structures nor UniParc sequences; for these, we performed gene prediction with Prodigal [41] on MAG assemblies and then predicted structures with AlphaFold3. This third category carries compound uncertainty: gene boundaries, protein sequences, and predicted structures are all derived computationally from metagenomic assemblies, but includes many phylogenetically important lineages (e.g., Baldrarchaeia, Hermodarchaeia, Wukongarchaeia) that are not otherwise represented in structural databases.
Structure prediction and quality assessment
Existing AlphaFold2 structures for the AFDB subset were retrieved from AFDB version 6. For the remaining 52,209 proteins (Prodigal and UniParc sources), multiple sequence alignments were generated using HHblits 3.3.0 [42] with 3 iterations against the UniRef30 2023_02 database [43] (E-value threshold 0.001). Structures were then predicted using AlphaFold3 v1 [18] with pre-computed MSAs (--norun_data_pipeline) on the TACC Lonestar6 computing cluster (NVIDIA A100 GPUs). Proteins were partitioned into tiers by length (Tier 0.5: 50–100 aa; Tier 1: 100–500 aa; Tier 2: 500–1000 aa; Tier 3: > 1000 aa) and processed in batches of 6 parallel GPU tasks per SLURM job. AF3 outputs were converted to AFDB-compatible mmCIF and PAE JSON formats for downstream processing.
Per-protein structure quality metrics were extracted from predicted structures for 116,275 proteins (93.7% of the dataset). Metrics included mean per-residue pLDDT confidence score (extracted from B-factor fields), fraction of disordered residues (pLDDT < 50), predicted TM-score, secondary structure composition (helix, sheet, and coil fractions), radius of gyration, and an overall quality category. The mean pLDDT score served as the primary indicator of structure prediction confidence and was used as a filter in the dark matter analysis. Predicted aligned errors (PAE) were generated for AlphaFold3 models but were not used as a classification-confidence threshold.
Domain assignment and annotation
Protein domains were identified using the Domain Parser for AlphaFold Models (DPAM) [19], which assigns ECOD [3] domain boundaries and classifications through iterative profile-profile search (HHsearch) [42], structural comparison (Foldseek, Dali) [22, 44], domain boundary refinement, and confidence scoring. DPAM was run against ECOD version 292 (June 2025). DPAM's primary classification output is a T-group (topology group) assignment, from which higher-level groupings, H-group (homologous superfamily) and X-group (possible homology), are derived by traversing the ECOD hierarchy. Each domain also receives a classification judge reflecting confidence: “good domain” (high-confidence ECOD match), “simple topology” (matched to a common fold with low specificity), “low confidence” (weak structural match), or “partial domain” (incomplete boundary). The “good domain” classification was used as the primary quality filter in downstream analyses.
DPAM was run on the LEDA HPC cluster using SLURM array parallelization. Separate processing runs accommodated the different input sources (AFDB structures, de novo AF3 predictions, and proteins requiring additional profile searches), but all used the same DPAM pipeline and reference data. Domain sequences were extracted from parent protein sequences using DPAM-assigned residue ranges.
Domain sequences were further annotated against Pfam-A version 38.2 using hmmscan from HMMER 3.1b2 [45] with the Pfam gathering threshold (GA) cutoff. Results were parsed from HMMER domain table output, extracting per-domain E-values, alignment coordinates, and posterior probabilities. In the standard ECOD accession workflow, T-group assignments from DPAM are combined with Pfam analysis to determine F-group (family) assignments [26].
Sequence and structural clustering
Clustering was performed at four levels to enable systematic comparison of sequence-based and structure-based groupings. Domain-level clustering used sequences and structures extracted at DPAM-assigned boundaries, while protein-level clustering used full-length sequences and structures.
Protein sequence clusters (PSC).
Full-length protein sequences were clustered with MMseqs2 [46] at 40% minimum sequence identity with 80% bidirectional coverage (--min-seq-id 0.4 -c 0.8 --cov-mode 0).
Domain sequence clusters (DSC).
Domain sequences were clustered with MMseqs2 using the same parameters as PSC (--min-seq-id 0.4 -c 0.8 --cov-mode 0).
Protein structural clusters (PXC).
Full-length protein structures were clustered with Foldseek (van Kempen et al., 2024) using a three-step pipeline: database creation (createdb), all-vs-all search, and greedy set-cover clustering (clust --cluster-mode 0). Search parameters were: TM-align alignment mode (--alignment-type 2), E-value threshold 0.001, minimum coverage 0.5, and no sequence identity filter (--min-seq-id 0.0).
Domain structural clusters (DXC).
Domain structures (extracted as individual mmCIF files) were clustered with Foldseek using the same pipeline and parameters as PXC.
Structural clustering parameters (E-value 0.001, 50% coverage, TM-align mode) were chosen to capture remote structural relationships while filtering spurious matches; TM-align mode scores full structural superposition rather than local fragment matching, appropriate for the domain-level comparisons central to this study.
Cross-clustering analysis
To quantify the extent to which structural clustering reveals relationships invisible to sequence comparison, we computed cross-cluster mappings between structural and sequence cluster dimensions. For each structural cluster, we identified the set of sequence clusters represented by its members. This analysis was performed at both protein (PXC to PSC) and domain (DXC to DSC) levels.
Characterization of the unclassified fraction
Proteins lacking any high-confidence DPAM domain assignment were systematically characterized through a series of filters designed to distinguish genuine structural novelty from explainable non-classification.
Proteins were first partitioned by DPAM classification status: “classified” (at least one domain with the “good domain” judge), “sub-threshold” (at least one DPAM domain of any judge category, but none reaching “good domain”), and “no domain” (no DPAM domain assignments at all). The no-domain category was further filtered to identify well-folded candidates for structural novelty. Proteins with a mean pLDDT below 70 were excluded as disordered or poorly predicted structures, and proteins shorter than 100 amino acids were excluded as below the effective sensitivity of template-based domain classification. The remaining proteins were assessed through their PXC cluster memberships: singletons (lacking any structural neighbor) were excluded, and proteins in clusters containing at least one member with a high-confidence domain assignment were classified as identifiable by transitivity. Proteins in clusters where at least one member carried any DPAM domain assignment (even sub-threshold) were classified as having partial signal. The residual, proteins in fully dark clusters where no member has any DPAM domain assignment at any confidence level, constitutes the irreducible core of structurally uncharacterized proteins.
Domain rescue by structural clustering
To assess whether structural clustering could validate sub-threshold DPAM classifications, we identified DXC clusters containing both high-confidence (“good domain”) and sub-threshold members. For each such mixed cluster, we determined the X-group consensus among the high-confidence members. Clusters where all high-confidence members shared the same X-group were designated as having consistent consensus. The X-group assignments of sub-threshold domains in consistent clusters were then compared against the cluster consensus and scored as agreeing, disagreeing, or receiving a new assignment (for domains lacking a prior X-group assignment). Protein-level rescue was defined as upgrading a protein from unclassified to classifiable when at least one of its sub-threshold domains was validated by a consistent rescue cluster.
Comparative analyses
Archaeal domain statistics were compared against our classification of Swiss-Prot [26] to contextualize classification coverage. Swiss-Prot protein lengths, domain classifications, and Pfam mappings were obtained from the UniProt and ECOD databases. ECOD hierarchy totals (X-groups, H-groups, T-groups) were computed from the full ECOD v292 classification. Superkingdom occupancy of H-groups (bacteria, eukaryota, archaea; Fig 3) was derived from the taxonomic annotations (NCBI Taxonomy) of classified domains in ECOD v292. Here “absent from bacteria” means no bacterial domain is currently classified into that H-group in ECOD v292: present classification coverage, not proof of biological absence.
Use of AI-assisted tools
Claude Opus 4.6 and Sonnet 4.6 (Anthropic) were used as a programming assistant during data analysis, figure generation, and manuscript preparation. Additionally, ChatGPT 5.3 was used for some manuscript preparation. All AI-assisted code was reviewed and validated by the authors. Claude was used for: database query construction and optimization, Python scripting for figure generation and data analysis, literature search assistance, and manuscript copyediting. All scientific interpretations, experimental design decisions, and domain expertise judgments were made by the authors. The AI tool did not generate scientific hypotheses, design experiments, or make classification decisions.
Supporting information
S1 Fig. AlphaFold2 and AlphaFold3 structure prediction quality comparison.
(A) Distribution of mean per-residue pLDDT scores for AF2 (AFDB-derived; n = 71,861; median 89.3) and AF3 (de novo predictions; n = 44,414; median 86.2). Dashed lines indicate medians. Both methods produce mostly high-confidence structures, with AF3 showing a slightly broader low confidence tail.
https://doi.org/10.1371/journal.pcbi.1014188.s001
(PNG)
S2 Fig. The archaeal fold repertoire is a conserved universal core across the six major groups.
Cladogram of the six major archaeal groups (schematic; topology after GTDB r220 and the two-domain archaeal phylogeny, the basal placement of DPANN and Deep-branching archaea remains debated), each tip annotated by its high-confidence (“good domain”) H-group repertoire. Stacked bars partition each group’s H-groups into the universal core present in all six groups (551 H-groups), folds shared by two to five groups, and group-specific folds present in a single group. Adjacent columns give H-group, X-group, and protein counts per group. The universal core is approximately constant across the phylogeny; repertoire size varies with lineage through shared accessory folds rather than new fold classes. Asgard archaea are the sister lineage to eukaryotes.
https://doi.org/10.1371/journal.pcbi.1014188.s002
(PNG)
S3 Fig. Sequence-based disorder confirms the unclassified fraction is not concealed structural novelty.
Intrinsic disorder predicted from sequence with metapredict V3 (threshold ≥ 0.5). (A) Distribution of per-protein predicted-disorder content for the 8,760 domain-free “dark” proteins versus 2,500 randomly sampled high-confidence classified proteins. Classified proteins are uniformly ordered; the dark fraction is bimodal, with a distinct distinct disordered population (≥ 50% of residues). (B) Mean pLDDT versus sequence-predicted disorder for the 5,854 domain-free proteins with structure-quality metrics; the weak correlation (r = -0.19) shows disorder and low prediction confidence are largely independent. Curated novel-fold candidates (n = 295) cluster at low disorder across the pLDDT range. Reference lines: pLDDT = 70 (dark-matter filter), disorder = 20% and 50%.
https://doi.org/10.1371/journal.pcbi.1014188.s003
(PNG)
S1 Table. Complete per-domain classification of the archaeal proteome dataset.
One row per classified domain (204,758 domains across 115,623 proteins; 36 columns): domain and protein identifiers (with cross-reference to the published accession and the EPP resource), source organism and GTDB taxonomy, structure provenance (pre-existing AFDB / de novo AlphaFold3) and mean pLDDT, the ECOD X-/H-/T-group assignment with DPAM confidence and match scores, Pfam annotation, and sequence and structural cluster membership at protein and domain levels (PSC, PXC, DSC, DXC). A full column dictionary accompanies the file. The same table is deposited at Zenodo (DOI 10.5281/zenodo.21656443).
https://doi.org/10.1371/journal.pcbi.1014188.s004
(GZ)
References
- 1. Cheng H, Schaeffer RD, Liao Y, Kinch LN, Pei J, Shi S, et al. ECOD: an evolutionary classification of protein domains. PLoS Comput Biol. 2014;10(12):e1003926. pmid:25474468
- 2. Schaeffer RD, Liao Y, Cheng H, Grishin NV. ECOD: new developments in the evolutionary classification of domains. Nucleic Acids Res. 2017;45(D1):D296–302. pmid:27899594
- 3. Schaeffer RD, Medvedev KE, Andreeva A, Chuguransky SR, Pinto BL, Zhang J, et al. ECOD: integrating classifications of protein domains from experimental and predicted structures. Nucleic Acids Res. 2025;53(D1):D411–8. pmid:39565196
- 4. Woese CR, Kandler O, Wheelis ML. Towards a natural system of organisms: proposal for the domains Archaea, Bacteria, and Eucarya. Proc Natl Acad Sci U S A. 1990;87(12):4576–9. pmid:2112744
- 5. Williams TA, Cox CJ, Foster PG, Szöllősi GJ, Embley TM. Phylogenomics provides robust support for a two-domains tree of life. Nat Ecol Evol. 2020;4(1):138–47. pmid:31819234
- 6. Varadi M, Bertoni D, Magana P, Paramval U, Pidruchna I, Radhakrishnan M, et al. AlphaFold protein structure database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 2024;52(D1):D368–75. pmid:37933859
- 7. Liu Y, Makarova KS, Huang W-C, Wolf YI, Nikolskaya AN, Zhang X, et al. Expanded diversity of Asgard archaea and their relationships with eukaryotes. Nature. 2021;593(7860):553–7. pmid:33911286
- 8. Spang A, Caceres EF, Ettema TJG. Genomic exploration of the diversity, ecology, and evolution of the archaeal domain of life. Science. 2017;357(6351):eaaf3883. pmid:28798101
- 9. Rinke C, Schwientek P, Sczyrba A, Ivanova NN, Anderson IJ, Cheng J-F, et al. Insights into the phylogeny and coding potential of microbial dark matter. Nature. 2013;499(7459):431–7. pmid:23851394
- 10. Castelle CJ, Wrighton KC, Thomas BC, Hug LA, Brown CT, Wilkins MJ, et al. Genomic expansion of domain archaea highlights roles for organisms from new phyla in anaerobic carbon cycling. Curr Biol. 2015;25(6):690–701. pmid:25702576
- 11. Cavicchioli R. Archaea--timeline of the third domain. Nat Rev Microbiol. 2011;9(1):51–61. pmid:21132019
- 12. Baker BJ, De Anda V, Seitz KW, Dombrowski N, Santoro AE, Lloyd KG. Diversity, ecology and evolution of Archaea. Nat Microbiol. 2020;5(7):887–900. pmid:32367054
- 13. DeLong EF. Archaea in coastal marine environments. Proc Natl Acad Sci U S A. 1992;89(12):5685–9. pmid:1608980
- 14. Gaci N, Borrel G, Tottey W, O’Toole PW, Brugère J-F. Archaea and the human gut: new beginning of an old story. World J Gastroenterol. 2014;20(43):16062–78. pmid:25473158
- 15. Bukhari SA, Caetano-Anollés G. Origin and evolution of protein fold designs inferred from phylogenomic analysis of CATH domain structures in proteomes. PLoS Comput Biol. 2013;9(3):e1003009. pmid:23555236
- 16. Karavaeva V, Sousa FL. Navigating the archaeal frontier: insights and projections from bioinformatic pipelines. Front Microbiol. 2024;15:1433224. pmid:39380680
- 17. Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–9. pmid:34265844
- 18. Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630(8016):493–500.
- 19. Zhang J, Schaeffer RD, Durham J, Cong Q, Grishin NV. DPAM: a domain parser for AlphaFold models. Protein Sci. 2023;32(2):e4548. pmid:36539305
- 20. Tunyasuvunakool K, Adler J, Wu Z, Green T, Zielinski M, Žídek A, et al. Highly accurate protein structure prediction for the human proteome. Nature. 2021;596(7873):590–6. pmid:34293799
- 21. Schaeffer RD, Zhang J, Medvedev KE, Kinch LN, Cong Q, Grishin NV. ECOD domain classification of 48 whole proteomes from AlphaFold structure database using DPAM2. PLoS Comput Biol. 2024;20(2):e1011586. pmid:38416793
- 22. van Kempen M, Kim SS, Tumescheit C, Mirdita M, Lee J, Gilchrist CLM, et al. Fast and accurate protein structure search with Foldseek. Nat Biotechnol. 2024;42(2):243–6. pmid:37156916
- 23. Rivera MC, Jain R, Moore JE, Lake JA. Genomic evidence for two functionally distinct gene classes. Proceedings of the National academy of sciences of the United States of America. 1998; 95(11):6239–6244.
- 24. Köstlbacher S, van Hooff JJE, Panagiotou K, Tamarit D, De Anda V, Appler KE, et al. Prediction of eukaryotic cellular complexity in Asgard archaea using structural modelling. Nat Microbiol. 2026;11(3):747–58. pmid:41786989
- 25. Leão P, Little ME, Appler KE, Sahaya D, Aguilar-Pine E, Currie K, et al. Asgard archaea defense systems and their roles in the origin of eukaryotic immunity. Nat Commun. 2024;15(1):6386. pmid:39085212
- 26. Schaeffer RD, Zhang J, Cong Q, Grishin NV (2025) ECOD: classification of domains in AFDB Swiss-Prot structure predictions.
- 27. Lotthammer J, Hernández-García J, Griffith D, Weijers D, Holehouse A, Emenecker R. Metapredict enables accurate disorder prediction across the Tree of Life. 2024.
- 28. Krishna SS, Majumdar I, Grishin NV. Structural classification of zinc fingers: survey and summary. Nucleic Acids Res. 2003;31(2):532–50. pmid:12527760
- 29. Liang J, Wang Z, He X, Li J, Zhou X, Deng Z. DNA modification by sulfur: analysis of the sequence recognition specificity surrounding the modification sites. Nucleic Acids Res. 2007;35(9):2944–54. pmid:17439960
- 30. Anantharaman V, Makarova KS, Burroughs AM, Koonin EV, Aravind L. Comprehensive analysis of the HEPN superfamily: identification of novel roles in intra-genomic conflicts, defense, pathogenesis and RNA processing. Biol Direct. 2013;8:15. pmid:23768067
- 31. Chen Y, Xiao H, Zhou J, Peng Z, Peng Y, Song J, et al. The in situ structure of T-series T1 reveals a conserved lambda-like tail tip. Viruses. 2025;17(3):351. pmid:40143278
- 32. Aravind L, Galperin MY, Koonin EV. The catalytic domain of the P-type ATPase has the haloacid dehalogenase fold. Trends Biochem Sci. 1998;23(4):127–9. pmid:9584613
- 33. Gomis-Rüth FX. Structural aspects of the metzincin clan of metalloendopeptidases. Mol Biotechnol. 2003;24(2):157–202. pmid:12746556
- 34. Chandonia J-M, Brenner SE. The impact of structural genomics: expectations and outcomes. Science. 2006;311(5759):347–51. pmid:16424331
- 35. Kihara D, Skolnick J. The PDB is a covering set of small protein structures. J Mol Biol. 2003;334(4):793–802. pmid:14636603
- 36. Levitt M. Nature of the protein universe. Proc Natl Acad Sci U S A. 2009;106(27):11079–84. pmid:19541617
- 37. Govindarajan S, Recabarren R, Goldstein RA. Estimating the total number of protein folds. Proteins. 1999;35(4):408–14. pmid:10382668
- 38. Wolf YI, Grishin NV, Koonin EV. Estimating the number of protein folds and families from complete genome data. J Mol Biol. 2000;299(4):897–905. pmid:10843846
- 39. Coulson AFW, Moult J. A unifold, mesofold, and superfold model of protein fold use. Proteins. 2002;46(1):61–71. pmid:11746703
- 40. Chaumeil P-A, Mussig AJ, Hugenholtz P, Parks DH (2022) GTDB-Tk v2: memory friendly classification with the Genome Taxonomy Database.
- 41. Hyatt D, Chen G-L, Locascio PF, Land ML, Larimer FW, Hauser LJ. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinform. 2010;11:119. pmid:20211023
- 42. Steinegger M, Meier M, Mirdita M, Vöhringer H, Haunsberger SJ, Söding J. HH-suite3 for fast remote homology detection and deep protein annotation. BMC Bioinform. 2019;20(1):473. pmid:31521110
- 43. Mirdita M, Driesch L, Galiez C, Martin MJ, Söding J, Steinegger M. Uniclust databases of clustered and deeply annotated protein sequences and alignments. Nucleic Acids Research. 2017; 45(D1):D170–D176.
- 44. Holm L. Using dali for protein structure comparison. Methods Mol Biol. 2020;2112:29–42. pmid:32006276
- 45. Eddy SR. Accelerated profile HMM searches. PLoS Comput Biol. 2011;7(10):e1002195. pmid:22039361
- 46. Steinegger M, Söding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol. 2017;35(11):1026–8. pmid:29035372