Fig 1.
Probabilities of shared alleles in pairwise sample comparisons for autosomal bi-allelic markers are derived from the list of genotype outcomes.
When no alleles are shared by descent (Z) (panel A, Z = 0), then the chance of seeing any specific combination of alleles is the product of the respective allele frequencies. When one (panel B, Z = 1) or both alleles (panel C, Z = 2) are shared by descent, then the possible number of genotype outcomes are reduced. The number of alleles identical by state (I) can be zero (panel A, lavender), one (all panels, no highlight), or two (all panels, green).
Table 1.
Calculation of P(I|Z) value for each marker.
Table 2.
Probabilities of different IBD and IBS states for different relationships.
Table 3.
Predicted HGMR values and standard deviations for different types of relationships assuming allele frequencies are evenly distributed between 0.1 and 0.9.
Table 4.
Comparison of the performances of the GRAF quadratic algorithm and KING 2.0 on finding identical pairs with different AGMR ranges.
Table 5.
Comparison of the performances of the GRAF quadratic algorithm and KING 2.0 on finding identical pairs with different numbers of SNPs with genotypes.
Table 6.
Running times and prediction accuracies of the sub-quadratic algorithm tested with datasets of different sample sizes and genotype missing rates.
Fig 2.
Distribution of all genotype mismatch rates of identical pairs detected by the quadratic algorithm and those reported by submitters.
All samples in the dbGaP Fingerprint Collection are compared. Types of submitter-reported relationships are color coded. Coral red: samples are reported to be from the same subjects; Purple: samples are from monozygotic twins; Gray: no relationship reported by submitters. Panel A shows the whole graph. Panel B shows the same graph, stretched on y-axis to show details of the bottom part of the graph.
Fig 3.
Distribution of homozygous genotype mismatch rates of sample pairs in four dbGaP studies.
Cyan curves show the distribution of HGMR values predicted with the assumption that the populations are homogeneous and random mating.
Fig 4.
Distribution homozygous genotype mismatch rates of sample pairs in the same four dbGaP studies as in Fig 2.
Graph is zoomed in to show the related pairs. Cyan curves show the distribution of HGMR values predicted with the assumption that the populations are homogeneous and random mating. Types of relationships reported by submitters in pedigree files are color coded. Red: parent/offspring; Blue: full sibling; Green: second degree relative; Yellow: third degree relative; Gray: no relationship reported by submitters.
Table 7.
Genotype missing rates and numbers of pairs of related samples reported in submitted pedigree files for four dbGaP studies.
Table 8.
Mean HGMR and AGMR values and correlation coefficients between HGMR and AGMR of all related subjects reported in the data files submitted to dbGaP.
Fig 5.
Distribution of both HGMR and AGMR values of all pairs of samples reported by submitters as closely related.
Each dot represents one pair of samples. Types of relationships reported by submitters are color coded. Purple: same subject or monozygotic twins; Red: parent/offspring; Blue: full sibling; Green: second degree relative; Brown: third degree relative. Each contour line shows the area that is predicted to include 95% of the sample pairs for each type of relationship.
Fig 6.
Comparison of GRAF and KING on determining subject relationships for four dbGaP studies.
Relationships self-reported in the pedigree files are color coded: red = parent-offspring; blue = full sibling; green = second degree; deep yellow = third degree. Cyan lines show the cutoff values to separate different types of relationships from one another.
Table 9.
δ values when different metrics are used to separate different types of relationships.
Kinship = kinship coefficient estimated by KING.
Fig 7.
An example of GRAF results displayed on dbGaP website for curators and data submitters to find discrepancies between genotypes and pedigree files submitted to dbGaP.
The graphs show GRAF results of one dbGaP study before the errors were corrected by the submitter. Relationships reported by the submitter are color coded. (A) Distribution of AGMR values. Coral red: Duplicate samples; Purple: monozygotic twins; Blue: first, second, or third degree relative; Gray: no relationship reported by submitter. (B) Distribution of HGMR values. Red: parent/offspring; Blue: full sibling; Green: second degree relative; Yellow: third degree relative; Grey: no relationship reported by submitter. (C) Distribution of both HGMR and AGMR values of pairs of related samples, excluding those from same subjects or monozygotic twins. Red: parent/offspring; Blue: full sibling; Green: second degree relative; Yellow: third degree relative; Gray: no relationship reported by submitter.