Table 1.
Summary of CNV detection in Chinese population.
Figure 1.
Copy Number Variable Region (CNVR) and Copy Number Variable Segment (CNVseg).
Due to the inconsistent calling of CNV boundaries, CNVR refers to a union of overlapping CNVs while a CNVseg represents the minimal units of CNV. In the two panels, black bars denote CNVR, below which each bar denote a CNV call in one individual. Red and blue bars represent deletion and duplication, respectively. (A). An example of CNVR and CNVsegs on chromosome 17 starting from 18,229,719 to 18,296,116. Dashed lines (green) dissect CNVR into small contiguous CNVsegs. (B). An example of a complex CNV region. The rectangle (orange) points out that one individual has both deletion and duplication in the region, to which we refer as complex CNVR in this study. (C). CNVR length distribution. Length distribution of 1440 CNVRs constructed in this study ranges from 1.03 kb to 1.73 Mb.
Figure 2.
Genomic CNV frequency distribution.
The height of bars indicates CNV frequency. Red and blue bars represent deletion and duplication respectively.
Figure 3.
Phylogenetic tree of Chinese ethnic groups constructed by UPGMA.
Phylogenetic tree of Chinese ethnic groups based on average pairwise genetic population distance between ethnic groups with 1,000 bootstrap replications by UPGMA.
Table 2.
Pairwise genetic distance among Chinese ethnic groups.
Figure 4.
Venn diagram (A) and bar chart (C) of CNV sharing results among African, Asian and European groups (each group with sample-size 425). Venn diagram (B) and bar chart (D) of CNV sharing results among Han Chinese, Chinese minority and Japanese groups (each group with sample-size 75). (E). Bar chart of sharing results in 708 non-singleton CNVRs (333 deletion CNVRs, 113 duplication CNVRs and 254 multi-allelic CNVRs) among 7 Chinese ethnic groups.
Table 3.
Pairwise CNVR sharing among different populations*.
Figure 5.
(A). FST distribution among world-wide populations. (B). FST distribution within Asian populations. CN: our 155 Chinese samples; TB: 41 Tibetan samples; HAN: our 80 Han Chinese samples.
Figure 6.
Population structure inferred by biallelic CNPs.
Principal component analysis was used to detect population structure. The results were plotted as the first principal component and second principal component. (A). Our Chinese samples with all HapMap groups, 349 CNPs. (B). Our Chinese samples with HapMap Asian groups, 606 CNPs. (C). Seven ethnic groups from our Chinese samples, 643 CNPs. (D). Han Chinese and Tibetans from our Chinese samples, 600 CNPs.
Figure 7.
Genomic population specific CNV significance plots.
The loss-allele and gain-allele frequency significance between Han Chinese and Tibetan is plotted as −log10 value in (A) and (B) respectively. Each point represents one CNV event. The dashed horizontal line indicates Bonferroni adjusted significance threshold.
Table 4.
Highly differentiated CNVRs between Han and other population.
Table 5.
Highly differentiated CNVRs between Tibetan and other populations.
Figure 8.
Linkage disequilibrium between CNPs and SNPs.
(A). SNPs taggability of all 155 Chinese, 80 Han and 41 Tibetan for common CNV regions (frequency>10%). SNPs located within 20 kb of either ends of CNV regions were included and SNP with maximum r2 among each CNVR was counted. (B). A deletion chr3:37,957,109–37,969,705 where LDs are almost perfect in Tibetan and CHD but weak in other populations. Black bar represents the location of this CNV.