Fig 1.
3D structure information for PDB entry 1STP.
A) Monomeric asymmetric unit. B) Unit cell containing asymmetric units within the crystal lattice.
Fig 2.
Quaternary structure annotations for PDB entry 1Z77.
A) Incorrectly annotated as monomer during PDB deposition. Correctly annotated as a dimer by B) PISA, C) EPPIC, and D) the primary publication.
Fig 3.
Overlapping structures between datasets.
Most of the duplicates occcurred between the Ponstingl and the Bahadur datasets. These two datasets share 89 protein structures, whereas only 2 structures were found to be common between the Bahadur and the Duarte datasets, and only 1 structure was detected in both the Ponstingl and the Duarte datasets.
Fig 4.
Sequence cluster construction and consistency score calculation.
A consistency score is calculated based on the distribution of stoichiometry and symmetry values within each cluster. (St: Stoichiometry, Sy: Symmetry, C(St, Sy): consistency score for given St and Sy).
Fig 5.
Cluster size vs. correct, incorrect, and N/A (not available cases) rates at 70% sequence identity.
Table 1.
Test set (20% of the original data) performances of the machine learning algorithms.
SVM outperforms BLR on a number of performance measures, whereas BLR has slightly better specificity and positive predictive value results.
Fig 6.
General workflow of the text mining approach.
After a publication upload or download, the term frequency-inverse document frequency (tf-idf) method is used to create a numerical data matrix from words, a hash table is created using the hash function, extracted sentences are classified using machine learning alorithms, remaining sentences are searched for experimental evidence and oligomeric state information is determined by using a majority rule.
Fig 7.
Overview of the PISA annotation procedure.
An XML file is created using PISA, the file is parsed to extract symmetry operators and corresponding chains and BioJava is used to assign stoichiometry and symmetry.
Fig 8.
Evaluation of the quaternary structure annotation of PDB entry 1Z77.
All methods (SC: sequence clustering, TM: text mining, PISA, EPPIC) are selected for the evaluation process. The output includes two parts: oligomeric state and symmetry. In the oligomeric state table, it states that 1Z77 (first column: PDB ID) has one oligomeric state prediction (second column: BA Number), and it is annotated as a monomeric structure in the PDB (third column: PDB). However, according to SC (fourth column: Sequence clustering), PISA (fifth column), EPPIC (sixth column), and TM (seventh column: Text Mining), 1Z77 is a dimeric protein. Therefore, our consensus result (eighth column) states that 1Z77 is a dimer. In the symmetry table one quaternary structure prediction is shown (second column: BA Number), having C1 symmetry (third column: PDB). However, according to SC (fourth columns: Sequence cluster), PISA (fifth column) and EPPIC (sixth column), 1Z77 has a C2 symmetry. Therefore our consensus result (seventh column) states that 1Z77 has a C2 symmetry. Red denotes divergence between current PDB annotation of oligomeric state and results provided by each of the four evaluation methods.
Fig 9.
Overall performance results for the benchmark dataset using four different methods and consensus approach (N/A: not available cases).
Fig 10.
Prediction agreement of the consensus approach.
A) Correct prediction agreements between methods. B) Incorrect prediction agreements between methods. C) Inconclusive prediction agreements between methods. D) Not-applicable predictions between methods.