Table 1.
AQ based on SISYPHUS_ID10.
Fig 1.
Comparison of DaliLite and CAB-align in terms of the agreement, reliability, and NormS values based on SISYPHUS_ID10.
(A) Scatter plots of agreement, (B) reliability, and (C) NormS. The numbers of pairs belonging to each area are indicated. For example, 557 CAD-align alignments had better agreement values than DaliLite, where the agreement values for CAB-align and DaliLite were both higher than 0.5. CAB-align, contact area-based alignment.
Fig 2.
Comparison of the AQ for six benchmark datasets.
(A) SCOPe_NR10_all (6,799 pairs), (B) SCOPe_NR10_e10 (3,660 pairs), (C) SCOPe_FAMILY_all (15,790 pairs), (D) SCOPe_FAMILY_e10 (5,730 pairs), (E) PDB30_e5 (182,907 pairs), and (F) PDB30_e10 (122,626 pairs). The methods are shown in order from top to bottom on the left (n = 0) of (A): CAB-align, DaliLite, FATCAT, and TM-align. CAB-align, contact area-based alignment; PDB, protein data bank.
Fig 3.
Consistency of alignments based on six datasets.
(A) SCOPe_NR10_all (7,384 triplets), (B) SCOPe_NR10_e10 (2,173 triplets), (C) SCOPe_FAMILY_all (50,630 triplets), (D) SCOPe_FAMILY_e10 (14,689 triplets), (E) PDB30_e5 (1,403,291 triplets), and (F) PDB30_e10 (790,623 triplets). PDB, protein data bank.
Fig 4.
ROC curves for the five alignment methods in the superfamily recognition test.
(A) NR10 benchmark dataset. (B) FAMILY benchmark dataset. (C) PDB30 benchmark dataset. PDB, protein data bank; ROC, receiver-operating characteristic.
Fig 5.
PRCs obtained using the NR10 and FAMILY datasets in the superfamily recognition test.
(A) NR10 benchmark dataset. (B) FAMILY benchmark dataset. (C) PDB30 benchmark dataset. PDB, protein data bank; PRC, precision-recall curve.
Table 2.
Classification performance with three benchmark datasets.
Fig 6.
ROC curves for the five alignment methods using the alignments returned by all methods.
(A) NR10 benchmark dataset. (B) FAMILY benchmark dataset. (C) PDB30 benchmark dataset. PDB, protein data bank; ROC, receiver-operating characteristic.
Fig 7.
PRC for the five alignment methods using the alignments returned by all methods.
(A) NR10 benchmark dataset. (B) FAMILY benchmark dataset. (C) PDB30 benchmark dataset. PDB, protein data bank; PRC, precision-recall curve.
Table 3.
Classification performance based on three benchmark datasets using the alignments returned by all methods.
Fig 8.
Flowchart illustrating the CAB-align procedure.
CAB-align comprises three main steps. Step 1: Rigid-body alignment method based on local structural similarity. Step 2: Rigid-body alignment method based on the global 3D structure superposition. Step 3: Iterative DP based on a modified CAD-score. CAB-align, contact area-based alignment; CAD, contact area difference; DP, dynamic programming.
Fig 9.
AQ of the components of CAB-align.
(A) SCOPe_NR10_all (6,799 pairs), (B) SCOPe_NR10_e10 (3,660 pairs), (C) SCOPe_FAMILY_all (15,790 pairs), and (D) SCOPe_FAMILY_e10 (5,730 pairs). CAB-align, contact area-based alignment.
Table 4.
AQ of the components of CAB-align based on SISYPHUS_ID10.
Table 5.
AUC and AUPRC scores for the components of CAB-align.
Table 6.
Computational time.
Fig 10.
Example of a structure comparison between two calmodulin-like proteins.
(A and C) Open-dumbbell conformation, 1ncx_A. (B and D) Closed conformation, 2sas_A. (A) Alignment of 1ncx_A by CAB-align. (B) Alignment of 2sas_A by CAB-align. (C) Alignment of 1ncx_A by DaliLite. (D) Alignment of 2sas_A by DaliLite. The aligned regions are rainbow color coded from blue to red. The calcium-binding regions assigned by UniProt are shown by sticks. CAB-align, contact area-based alignment.
Fig 11.
Examples of structural alignments between two calmodulin-like proteins.
(A) Structural alignment by CAB-align. (B) Alignment graph produced by HHalign and CAB-align. (C) Structural alignment by DaliLite. (D) Alignment graph produced by HHalign and DaliLite. (A and C) The asterisks represent the calcium-binding regions. CAB-align, contact area-based alignment.
Fig 12.
Example of a structure comparison using the SISYPHUS benchmark dataset.
(A, C) 2c2f_A. (B, D) 1j30_A. (A) Alignment of 2c2f_A by CAB-align. (B) Alignment of 1j30_A by CAB-align. (C) Alignment of 2c2f_A by DaliLite. (D) Alignment of 1j30_A by DaliLite. The aligned regions are rainbow color-coded from blue to red. CAB-align, contact area-based alignment.
Fig 13.
Examples of structural alignments between 2c2f_A and 1j30_A.
(A) Alignment graph produced using SISYPHUS and CAB-align. (B) Alignment graph produced using SISYPHUS and DaliLite. CAB-align, contact area-based alignment.
Fig 14.
Calculation of the residue–residue contact area matrix for the input PDB file (PDBID: 1dqx_A).
PDB, protein data bank.
Fig 15.
Calculation of the similarity matrix M from the given alignment.
(A) The given alignment. The black cells represent the aligned pairs in protein A and B. (B) In the matrix M, the similarity score of the residue pair (ka,kb) is calculated from the comparison between other aligned positions. The gray cells are ignored. The curved lines represent the comparison between two residue pairs (Eq 12).
Fig 16.
Calculation of the similarity matrix M′ with a window size of three.
(A) In the matrix M′, the similarity score for the residue pair (ka,kb) is calculated from the two orange cells (M(ka–1,kb–1) and M(ka+1,kb+1)) and the red cell M(ka,kb). (B) Similarity score for the residue pair (ka–1,kb–1). (C) Similarity score for the residue pair (ka,kb). (D) Similarity score for the residue pair (ka+1,kb+1). The black cells represent the aligned positions in the given alignment. The gray cells represent the ignored pairs.
Fig 17.
Protocols employed for iterative DP with a window size of three.
(A) The given alignment. (B) The similarity score matrix M′ is updated by the alignment. (C) DP is performed using the M′ and a new alignment is generated. DP, dynamic programming.
Fig 18.
Distributions of the normalized score, NormS.
The contour lines are plotted at interval values of 10.0 for NormS. The vertical line represents the average rate of S (i.e., 0.5*(S/TA+S/TB)). The horizontal line represents the similarity score S. (A) TA:TB = 1:1, (B) TA:TB = 1:2, and (C) TA:TB = 1:5. NormS, normalized similarity score.