Figure 1.
Comparison of comparative modeling, BCL::Fold and Rosetta.
In comparative modeling the protein backbone is constructed partially from a target-template alignment, followed by loop construction and side chain building. On the other hand, de novo methods, such as Rosetta, only take advantage of the decoupling of backbone placement and the side chain building. BCL::Fold also decouples the construction of loops from assembly of secondary structure elements, similar to comparative modeling. Although disconnecting these steps makes computation more feasible by splitting the total search space into manageable portions, they are not absolute and in order to address these issues SSE placement has to be refined before loop building and side chain construction. While side chain conformations are not explicitly added to the model in the early stages of de novo structure prediction, they are implicitly represented throughout the process through knowledge-based potentials.
Figure 2.
(A) Generation of secondary structure element (SSE) pool. The secondary structure prediction methods, PSIPRED and JUFO, have been combined to achieve a consensus three state secondary structure prediction. For a given amino acid sequence, stretches of sequence with consecutive α-helix or β-strand predictions above a given probability threshold are identified as α-helical and β-strand SSEs and added to the pool of SSEs to be used in the assembly protocol. (B) Assembly of SSEs. The initial model only has a randomly picked SSE from the SSE pool. At each single iteration step, a move is picked randomly and applied to produce a new model. The details regarding utilized moves are given in the next panel. (C) Energy evaluation using knowledge based potentials. After each change, the model is evaluated using knowledge based potentials. These include loop, loop closure, amino acid environment, amino acid pair distance, amino acid clash, SSE packing, strand pairing, SSE clash, contact order and radius of gyration. (D) Monte Carlo Metropolis minimization. Based on the energy evaluation, models with lower energies than the previous model are accepted, while models with higher energy can be either accepted or rejected based on Metropolis criteria. The accepted models are further optimized, in case of rejected models, the minimization continues with the last accepted model. The minimization is terminated after either a specified total number of steps or a specified number of rejected steps in a row. The protocol consists of two such minimizations, one for assembly and one for refinement.
Table 1.
Benchmark set of proteins.
Table 2.
Secondary structure pool statistics for the benchmark proteins.
Figure 3.
SSE-based moves allow rapid sampling in conformational search space.
The types of moves used in the BCL::Fold protocol are explained with a representative move set. (A) Single SSE moves: These moves can include adding a new SSE to the model from the pool as well as translations/rotations/transformations. (B) SSE pair moves: One of the SSEs in the pair can be removed, the locations can be swapped and one can be rotated around the other SSE which is used as a hinge to define rotation axis. (C) Domain based moves: These moves act on a collection of SSEs such as a helical domain or β-sheet. The examples show how the locations of strands can be shuffled within a β-sheet or how an entire β-sheet can be flipped or translated.
Figure 4.
Correlation of moves used in BCL::Fold.
The correlation of all move pairs is depicted as heat maps for (A) assembly moves (B) refinement moves. For both heat maps, the moves are ordered in the same order as in Tables S1 and S2 respectively.
Figure 5.
Structures for a selection of best RMSD100 complete models generated by BCL::Fold.
Best complete models by RMSD100 with a predicted pool generated by BCL::Fold for a selection of proteins. The generated models are rainbow colored and superimposed with the native structure (gray) for the following proteins. The numbers refer to the RMSD100 of the models: (A) 1GYUA –6.39Å (B) 1ICXA –6.46Å (C) 1ULRA –4.73Å (D) 1X91A –4.49Å (E) 1J27A –3.72Å (F) 1TP6A –6.83Å (G) 2CWRA −7.61Å (H) 2RB8A –5.09Å (I) 1RJ1A –5.33Å (J) 1TQGA 2.44Å (K) 2HUJA –3.37Å (L) 3OIZA –6.21Å (M) 2V75A –3.55Å.
Figure 6.
Structures for a selection of best RMSD100 SSE-only models generated by BCL::Fold.
Best SSE-only models by RMSD100 with a predicted pool generated by BCL::Fold for a selection of proteins. The generated models are rainbow colored and superimposed with the native structure (gray) for following proteins. The numbers refer to the RMSD100 of the models: (A) 1GYUA –4.11Å (B) 1ICXA –6.07Å (C) 1ULRA –3.61Å (D) 1X91A –4.08Å (E) 1J27A –3.15Å (F) 1TP6A –5.74Å (G) 2CWRA −5.66 Å (H) 2RB8A –2.91Å (I) 1RJ1A –4.86Å (J) 1TQGA –1.92Å (K) 2HUJA –2.57Å (L) 3OIZA –5.76Å (M) 2V75A –3.11Å.
Table 3.
Best RMSD100 and CR values for models generated by BCL and Rosetta.
Figure 7.
Comparison of best RMSD100 and CR values for BCL and Rosetta.
Scatter plot comparing (A) best RMSD100 or (B) best CR SSE-only (left) and complete (right) BCL models vs. Rosetta models. The BCL models considered are from BCL::Fold runs using predicted SSE pools. (B) Scatter plot comparing best CR SSE-only (left) and complete (right) BCL models vs. Rosetta models. The BCL models considered are from BCL::Fold runs using predicted SSE pools.
Figure 8.
Determinants of high CR values in BCL and Rosetta models.
(A) Plot of sequence length vs. relative contact order (RCO) for all benchmark proteins. (B) Plot of percentage of amino acids found in SSEs vs. maximum Q3 value achieved from JUFO or PSIPRED pools for all benchmark proteins. Individual plots are presented for models from BCL::Fold runs using predicted SSE pools (left panels) and Rosetta models (right panels). Points in both (A) and (B) are colored according to the best CR value achieved for that benchmark protein in BCL runs using predicted SSE pools and complete models;<20% (red), 20% to 40% (orange), 40% to 60% (green) and>60% (dark green).
Figure 9.
The best submitted model out of the 5 top submissions by RMSD (rainbow colored) superimposed with the native structure for (A) T0608_1–89 residues, 4.3Å RMSD (B) T0580 - 105 residues 4.44Å RMSD, (C) T0619 - 111 residues, 5.86Å RMSD (D) T0602 - 123 residues, 7.75Å RMSD (E) T0630 - 132 residues, 8.42Å RMSD (F) T0627 - 261 residues, 8.90Å RMSD.