Figure 1.
Analysis of sorghum CS using DArTs, RFLPs and SSRs.
(A) Average log-likelihood and standard errors obtained with STRUCTURE software. For the three marker systems, the log-likelihood reached a plateau around K = 6. (B) The ΔK parameter [69] enabled the identification of high values at K = 2 (DArTs, RFLPs) and K = 4 (DArTs, SSRs), and lower peaks at K = 7 (DArTs, SSRs) and K = 6 (RFLPs, although this peak was not visible with the Y axis-scale used). (C) Proportion of unassigned accessions at the 0.6 membership threshold for each marker system. This proportion varied with the marker systems and the number of groups, K. To assess the stability of accession assignments to the different genetic groups between runs for the different marker systems, the average dissimilarity indices between runs for all accessions were computed and are reported in (D) according to calculations presented in the Material and Methods section. This analysis revealed that SSR markers provided the least stable assignment across runs for the same K level. Lastly, assignment stability across marker systems is reported in (E) through the calculation of the average dissimilarity index for all accessions for each pair of marker systems (see text for details). This analysis indicated that the assignments obtained with the SSRs were the most divergent from the other marker systems (DArTs and RFLPs).
Table 1.
Genetic diversity parameters within the Core Sample (171 accessions) and within the 6 genetic groups identified.
Figure 2.
Sequential identification of the genetic groups through model-based analysis as revealed by the different marker systems and comparison of the most relevant model-based structure (K = 6) with distance-based method analysis.
(A) Genome composition of accessions for different levels of structure. Each sample is represented in K dimensions, with K being the number of hypothetical genetic groups that compose the collection (K ranging from 2 to 10). Three different datasets were tested: 713 DArTs, 60 RFLPs, and 40 SSRs. Each accession on the X-axis is represented by K colours (each corresponding to a genetic group) ordered according to a decreasing genome fraction on the Y-axis. For each dataset, 171 accessions were ordered according to DArT assignments, with a decreasing proportion of genome assigned to the main groups according to STRUCTURE at K = 10. (B) Neighbour-Joining tree of the CS with colour projection of the six groups obtained with the model-based method at the 0.6 membership threshold at K = 6. The Neighbour-Joining tree is based on the genetic similarities between accessions calculated as the proportion of shared alleles of the DArT markers (Sokal and Michener modality index). The colours correspond to the genetic groups obtained at K = 6 in (A). Accessions assigned at a proportion <0.6 are coloured in grey. Group A in pink includes D and B from India, C and CB from China. Group B in blue includes C and D from Africa. Group C in red includes G from Western Africa. Group D in spring green includes Gm from Western Africa. Group E in orange includes G from Southern Africa and Asia. Group F in green includes K and KC from Southern Africa. Clusters were identified by Deu et al. (2006).
Table 2.
Pairwise differentiation (Fst) between the six genetic groups as obtained with three different marker systems (DArT, RFLP, SSR).
Figure 3.
Ability of the marker systems to describe the genetic structure of the sorghum Core Sample.
Total, intra- and intergroup accession dissimilarity correlations obtained with an increasing number of the three different marker systems (DArT, RFLP, SSR) were computed according to the data resolution statistic developed by Van Hintum [31]. An enlargement of the graph for marker numbers comprised between 0 and 150 is provided at the bottom right-hand side of the figure. For the whole CS, the description of diversity can only be considered saturated for DArT markers. An analysis of the dissimilarity correlations at the intra- and intergroup levels indicated that SSR markers were the least efficient in describing the intergroup structure.
Figure 4.
Evolution of linkage disequilibrium for different sample sizes and distances.
Mean r2 and the proportion of significant pairwise r2 (i.e greater than P95) were computed for the CS and for subsets of accessions ranging from 25 to 150. Accessions were sampled using two strategies. Firstly, random samples of 25, 50, 75, 100, 125 and 150 genotypes were extracted to calculate LD statistics. For each sample size, 10 random samples were analysed in order to estimate standard errors of the estimations (A). Secondly, a procedure designed to define subsets of genotypes minimizing their redundancy and limiting the loss of diversity was used (MLST reported in (B)). A comparison of these two sampling approaches for three sample sizes is provided in (C) and indicates that, with both strategies, mean r2 was overestimated for all distance classes in small samples (n = 25). The proportion of significant pairwise r2 was always higher with the MLST, compared to the random approach, especially for small distances and small sample sizes, highlighting the efficiency of this algorithm in providing a more accurate local LD estimation through the reduction of background LD. After correcting for background structure, mean r2 decreased with distance from 0.2 (for the 0–10 kb distance class) to 0.02 and stabilized after 100 kb (B). Contrary to random sampling, an increase in mean r2 (B) was observed with the MLST approach when sample sizes greater than 100 accessions were considered, suggesting that redundancy, and thus background structure, was introduced after this sample size. These results indicate that, for the CS, a sample size of 100 accessions carefully selected to avoid redundancy and maximize diversity would be an optimized sample size for LD estimation.
Table 3.
Evolution of linkage disequilibrium with physical distance in the Core Sample.
Table 4.
Number of markers required for association mapping studies.
Table 5.
DArT markers presenting evidence of selection.
Figure 5.
Concomitant variability of the photoperiod sensitivity index with the allelic frequency of the DArT SbMITE-188058.
The data for the photoperiod sensitivity index (Kp, which corresponds to the decrease in the duration of the vegetative phase between two sowing dates) presented in (A) were obtained from Clerget et al. [70] who used the same collection of accessions. Kp varies from 0 for photoperiod-insensitive varieties, which do not change the duration of their vegetative phase with the sowing date, to 1.0 for the most strongly photoperiod-sensitive varieties which maintain their calendar date of flowering constant by reducing the duration of their vegetative phase. A total of 136 accessions with a membership coefficient greater than the 0.6 threshold were considered. These accessions corresponded to 40 accessions from group A, 27 from group B, 21 from group C, 11 from group D, 17 from group E and 22 from group F. An Anova analysis indicated highly significant differences between the genetic groups (p value<2.2e-16) with genetic group F harbouring a low photoperiod sensitivity index. An analysis of the variability in the allelic frequency of the DArT marker SbMITE-188058 (FI847787) located at 11.7 kb from a homologue of the FAR-RED IMPAIRED RESPONSE 1 protein isolated in Arabidopis thaliana (B) highlighted the specificity of group F, suggesting a potential role of this gene in the genetic control of variability in photoperiod sensitivity.