Figure 1.
Breakdown of proteins by organism in the Riley dataset [8].
More than half of the organisms have only two proteins each. The top four organisms with the largest number of proteins are Fruitfly, Yeast, Worm and Human.
Table 1.
A sample of organisms in the Riley dataset with their respective protein, domain and gold domain sizes.
Figure 2.
Domains do not occur with the same frequency.
There are rare and promiscuous domains (left). Rare domains are domains which appear infrequently in the set of proteins. The rare domains outnumber the promiscuous domains several fold. The rare domains occur less frequently in the set of gold domains than in the set of all domains. In general, the more frequently a domain occurs, the more likely it is to be a gold domain (right). The same patterns are observed when proteins are confined to single organisms (Figure S1 in File S1).
Figure 3.
Ratio of promiscuous to rare domains.
At rare domain threshold x, domains which occur at most x times in a protein set, i.e. N(d) ≤ x, are classified as rare. When the rare domain threshold is 1, 58.5% of all domains are rare, and the remaining 41.5% are promiscuous, i.e. occurs 2 or more times in the set of proteins. When the rare domain threshold is 2, 80.4% of all domains are rare. The percentage of domains that are promiscuous drops to 6.1% when the rare domain threshold is 5.
Figure 4.
Breakdown of proteins by architecture (domain composition).
As more domains are classified rare, the proportion of proteins comprising only-rare domains increases, while the proportion of proteins comprising only-promiscuous domains decreases. Mixed architecture proteins comprise both rare and promiscuous domains. They make up the largest percentage (at least a third) of the protein population when the rare domain threshold is between 2 and 16. The peak occurs at rare domain threshold 4 where 37.8% of the proteins have mixed architecture. Note that the concept-based scoring method proposed in this paper does not require or depend on the specification of a rare domain threshold. The purpose of the analysis is this figure is to provide prima facie evidence for the feasibility of the concept-based scoring method which relies on a strong presence by mixed architecture proteins for a successful outcome.
Figure 5.
Domain coverage by protein architecture type.
Coverage of all domains by protein architecture type (top). “onlyrare” refer to proteins comprising one or more rare domains. “onlypromis” refer to proteins comprising one or more promiscuous domains. “mixed” refer to proteins comprising rare and promiscuous domains. Coverage of the 642 gold domains by protein architecture types (bottom). A gold domain is a domain that participates in at least one gold standard domain-domain interaction (GDDI). For rare domain threshold values between 2 and 4 inclusive, even though only-rare proteins involve a larger fraction of all domains, they cover a smaller fraction of gold domains than only-promiscuous proteins. Thus gold domains are more likely to be promiscuous domains. Mixed architecture proteins provide the largest coverage of gold domains when the rare domain threshold is between 3 and 5 inclusive. This range lies within the range where mixed architecture proteins make up the largest proportion of the protein population (Figure 4). So there exists a sweet spot where mixed architecture proteins are the most popular protein type and provide the largest coverage of gold domains.
Figure 6.
The OA (fully-labeled) concept lattice for the S. pombe context in Table 2.
Table 2.
A cross-table representing the relation between proteins and domains associated with S. pombe in the Riley dataset.
Figure 7.
The oa (reduced) concept lattice for the context in Table 2.
The OA concept lattice was presented in Figure 6. The oA and Oa concept lattices are given in Figures S2 and S3 in File S1, respectively.
Figure 8.
Key steps in the proposed concept-based scoring and ranking method for domain-pairs.
See text for further explanation and Figures S6 and S9 in File S1 for a walk-through on how to compute the <CB, PG> value for a domain-pair.
Figure 9.
Concept-pair promiscuity decreases as score increases.
Spearman’s rank correlation rho for OA is −0.4934008, and for oA is −0.3893343. Promiscuity of a domain-pair (a, b) is [N(a)+N(b)]/2 where N(d) is the number of times domain d occurs in a set of proteins. For a concept-pair (ci, cj), promiscuity is the average promiscuity of all domain-pairs in AL(ci) × AL(cj), and the score is the log-odds ratio of interacting protein-pairs to non-interacting protein-pairs in OL(ci) × OL(cj). For OA concept-pairs that produce a score in [−12, −10), the median promiscuity is 133.5 and the mean promiscuity is 129.2; when the score is 0.0, the median promiscuity falls to 13.25 and the mean promiscuity is 23.47. There were no oA scores smaller than −12.0. In non-attribute-reduced concept lattices (OA and oA), a domain-pair can be generated by more than one concept-pair (ci, cj) through the cross-product of their attribute-label sets, i.e. AL(ci) × AL(cj). This paves the way for a more promiscuous domain-pair to have the same score as a less promiscuous domain-pair. The piggy-backing mechanism takes effect when a domain-pair improves its score because either one or both of its domain partners happen to occupy the same attribute-label set as one or more rarer domains.
Figure 10.
Impact of domain shuffling on domain architecture (top).
Each point plots the minimum and maximum domain frequency in a protein. For example, the original domain set for protein 949, D(949) = {PBD, Pfam-B_2441, Pkinase} (Table 2). The minimum and maximum occurrence values for D(949) are 1 and 5 respectively. A point (x, y) in the plot denotes the minimum and maximum domain occurrence in a protein. Prior to domain shuffling, there are proteins with both rare and promiscuous domains, as shown by the black markings in the upper left of the plot. After domain shuffling, proteins have domains which are either rare only or promiscuous only, as shown by the orange markings on the y = x line. Domain shuffling changes the original heterogeneous domain architecture to a homogeneous one in terms of domain occurrence. Gold domain coverage by protein type after domain shuffling (bottom). At all rare domain threshold values, at least 99.5% of the 642 gold domains are covered by either proteins comprising only rare domains or proteins comprising only promiscuous domains. When the rare domain threshold is 4, the coverage comes very close to a 50∶50 split. Compare with Figure 5 (bottom) for gold domain coverage by protein type before domain shuffling.
Figure 11.
Scatter-plot of GDDI rank vs. promiscuity, scenario A Pe = 1.0.
All the 177,233 putative DDIs were ranked as described in the text, and the ranks of GDDIs were extracted to create the plots. Promiscuity of a domain pair (a, b) = [N(a)+N(b)]/2 where N(d) is the number of times domain d occurs in a protein set. Only the concept lattices which are not attribute-reduced (OA and oA) exhibit the desired negative relationship, which means they tend to rank promiscuous GDDIs more highly. The relationship is strongly positive when the Oa rankings are used. Oa results are identical to the Associative method which is known to penalize promiscuous domain-pairs. There is also a tendency for the oa concept lattice to rank promiscuous GDDIs less highly, but this positive relationship is not so apparent because scoreless GDDIs are included in the plot (they start at rank 49,378 and onwards to the right of the plot). When the oa concept lattice is used, only 350 of the 783 GDDIs have CB scores; the remaining GDDIs are scoreless and are ranked randomly but below the GDDIs with CB scores. The same pattern of relationships is observed with evaluation scenarios B and C (Figures S11 and S12 in File S1).
Figure 12.
Scatter-plot of GDDI rank vs. promiscuity, scenario D shuffled domains, Pe = 1.0.
Mixed architecture proteins play a critical role in the ability of non-attribute-reduced concept lattices to rank promiscuous GDDIs highly. All the 194,752 putative DDIs were ranked as described in the text, and the ranks of GDDIs were extracted to create the plots. Promiscuity of a domain pair (a, b) = [N(a)+N(b)]/2 where N(d) is the number of times domain d occurs in a protein set. The concept lattices which are not attribute-reduced (OA and oA) no longer exhibit the desired negative relationship (Figure 11). Instead, they tend to rank promiscuous GDDIs less highly. The relationship is still strongly positive when the Oa rankings are used. Oa results are identical to the Associative method which is known to penalize promiscuous domain-pairs. There is still also a tendency for the oa concept lattice to rank promiscuous GDDIs less highly, but this positive relationship is not so apparent because scoreless GDDIs are included in the plot (they start at rank 28,457 and onwards to the right of the plot). When the oa concept lattice is used, only 59 of the 214 GDDIs have CB scores; the remaining GDDIs are scoreless and are ranked randomly but below the GDDIs with CB scores.
Figure 13.
Effect of domain-shuffling on the relation between attribute-label frequency and domain frequency.
The scatter-plot at the bottom zooms in on the first 100 domain frequency values. There is a strong positive correlation prior to domain-shuffling (scenario A) which is lost after domain-shuffling (scenario D). The lost of this strong positive correlation impairs piggy-backing (Figure 14).
Figure 14.
Effect of domain-shuffling on PG values and piggy-backing.
In both scenarios A and D, OA produces a larger range of PG values than oA (top). Prior to domain-shuffling (A), the range of PG values is [0, 523] for OA and [0,21] for oA. Post domain-shuffling (D), the range of PG values is [0, 16] for OA and [0, 9] for oA. A reason for this is OA uses many more concepts than oA to score domain-pairs (Table S4 in File S1). In both scenarios A and D, OA has more piggy-backing, as evidenced by its much higher proportion of domain-pairs with PG >0. A domain-pair with PG >0 means its CB score is the result of one or more piggy-backs. The text discusses the pros and cons of OA’s higher piggy-backing potential. For oA, the proportion of domain-pairs with PG >0 drops from 4.83% (8555/177233) to 1.69% (3284/194752) when the domains are shuffled. With fewer piggy-backs, the results for oA deteriorate. PG is always 0 for attribute-reduced concept lattices (oa and Oa).
Figure 15.
Recovery of GDDIs [6].
GDDIs make up 0.44% of DDIs in A, 0.60% in B, 0.21% in C and 0.11% in D. Median promiscuity of GDDIs to DDIs is 7.0∶5.0 in A, 8.5∶5.5 in B, 9.5∶5.5 in C and 15.0∶4.0 in D. Except for the concocted scenario D, at high Specificity (FPR ≤ 0.2), concept-based rankings produced with concept lattices that are not attribute-reduced (OA and oA) outperform (have higher Sensitivity or larger TPR) those by concept lattices that are attribute-reduced (Oa and oa). In scenarios A to C, oA outperforms OA. This outcome holds even when DOMINE [18] and 3did [19] domain-pairs are used to evaluate the rankings (Figure 16). The quality of the oA domain-pair rankings becomes more evident by examining its top 100 domain-pairs (Figure 17 and File S2). OA’s poorer performance in scenarios A to C is attributed to “excessive” piggy-backing. Scenario D shows that oA’s GDDI recovery is more sensitive to changes in protein domain architecture than OA’s. The changes introduced by domain shuffling in D reduce oA’s piggy-backing potential and its GDDI recovery suffers as a result. In contrast, OA is able to retain some of its previously “excessive” piggy-backing potential (Figure 14). Nonetheless, the GDDIs recovered at low FPR by OA in scenario D are of the less promiscuous variety (Figure 12). In all four scenarios, the oa concept lattice leaves at least 72% domain-pairs and at least 55.30% GDDIs scoreless. Also, the Oa rankings produced the worst GDDI recovery performance in all four scenarios.
Figure 16.
Recovery of domain-pairs in DOMINE [18] (top), and in 3did (Jul-25-2013 release) [19] (bottom).
At high Specificity (FPR ≤ 0.2), oA’s TPR dominates the other three rankings. All rankings were made under scenario A conditions.
Figure 17.
Number of GDDIs [6], DOMINE domain-pairs [18] and 3did domain-pairs [19] found in the top 100 domain-pairs.
Of the four rankings (all made under scenario A conditions), oA contains the most number of GDDIs, DOMINE domain-pairs and 3did domain-pairs in the top 100 domain-pairs (File S2). The 26 GDDIs for oA intersect with, but is not the same set as, the 26 3did domain-pairs for oA. Amongst oA’s top 100 domain-pairs is domain-pair (TPR, WD40) which is not a GDDI and whose interaction is neither recorded in DOMINE nor predicted by 3did. However, a possible 3D structure for the obesity-related protein adipose (adp) involves interaction between TPR and WD40 domains [20].
Figure 18.
Nye test pass rate for the four concept lattice types in the four test scenarios.
A higher pass rate means a larger proportion of GPPIs have a GDDI as the highest ranking DDI. Except for D, concept lattices which are not attribute-reduced (oA and OA) have significantly higher pass rates than concept lattices that are attribute-reduced. The number of GPPIs is different in each scenario since GPPIs depend on GDDIs and PPIs, both of which are affected in scenarios B to D (Table S3 in File S1). GPPIs with only one DDI are included in the counts. The 2,326 GPPIs in A includes 546 single-DDI GPPIs. The Nye test results supports the hypothesis that in addition to a concept lattice that is not attribute reduced, mixed architecture proteins are also necessary to create favourable conditions for concept-based scoring to do well.
Figure 19.
Nye test results by difficulty.
The Nye test is more difficult as the number of DDIs per GPPI increases (Figure S13 in File S1). Each bar shows the fraction of GPPIs with x number of DDIs that pass the Nye test, i.e. where a GDDI is the highest ranking DDI, in scenario A. As the level of difficulty increases, concept-based scoring and ranking made with the attribute-reduced concept lattices (Oa and oa) become less able to pass the Nye test. With a few exceptions, regardless of difficulty, more GPPIs pass the Nye test with the OA ranking than with the oA ranking. This is attributable to OA’s greater potential for piggy-backing (Figure 14). However, the piggy-backing mechanism is also available to promiscuous non-GDDIs, and this causes some GPPIs to fail the Nye test when using the OA ranking. For example, the GPPI with 50 DDIs is (2252, 2530). This GPPI is supported by only one GDDI (AAA, AAA), which has a promiscuity of 100. oA ranks (AAA, AAA) as the highest domain-pair for GPPI (2252, 2530) and so the Nye test is passed. OA ranks (AAA, AAA) as the second highest, below (Pfam-B_1, AAA) which has a promiscuity of 189.5 but is not a GDDI, and so the Nye test is failed. The Nye test is also passed by oa since (AAA, AAA) is the only domain-pair with a score for GPPI (2252, 2530). However, oa’s Nye test performance on a GPPI with this high level of difficulty is more the exception than the norm.
Figure 20.
Concept-based scoring and ranking in the absence of Pfam-B domains.
Rankings were made under scenario A conditions. Recovery of 3did domain-pairs for yeast (top-left) and human (bottom-left). Without Pfam-B domains, the ROC curve for Oa is no longer below the y = x line as it was in Figure 15. Nonetheless, at high Specificity (FPR ≤ 0.2), the Sensitivity (TPR) of non-attribute-reduced concept lattices (OA and oA) still dominate that of the attribute-reduced concept lattices (Oa and oa) although now it is no longer as clear as it was in Figure 15 that oA’s TPR dominates OA’s TPR. The point is not to quibble about the difference between OA and oA, but between attribute-reduced where piggy-backs are impossible and non-attribute-reduced where piggy-backs are possible. The ROC curves show that in the absence of Pfam-B domains and on a more current dataset for single organisms, both OA and oA still outperform both Oa and oa. The Nye test for yeast (top-right) and for human (bottom-right) is also more successfully passed by both OA and oA than either Oa or oa.
Figure 21.
Number of 3did domain-pairs [19] found in the top 100 domain-pairs for Yeast and for Human.
Without Pfam-B domains, both oA and OA still have more 3did domain-pairs in their top 100 than both Oa and oa. However, in contrast to Figure 17, now OA has more 3did domain-pairs than oA. This is because there are no promiscuous Pfam-B domains to cloud OA’s ranking.