Detecting the number of clusters of individuals using the software <scp>structure</scp>: a simulation study
2005-05-11
SCID: 54.1/vqvmt3my
Discuss with AI
AFLP vs. microsatelliteDeltaK statisticSTRUCTURE softwareindividual-based simulationnumber of clusters (K)
Figures from the paper
Abstract (AI)
The identification of genetically homogeneous groups of individuals is a long standing issue in population genetics. A recent Bayesian algorithm implemented in the software STRUCTURE allows the identification of such groups. However, the ability of this algorithm to detect the true number of clusters (K) in a sample of individuals when patterns of dispersal among populations are not homogeneous has not been tested. The goal of this study is to carry out such tests, using various dispersal scenarios from data generated with an individual-based model. We found that in most cases the estimated 'log probability of data' does not provide a correct estimation of the number of clusters, K. However, using an ad hoc statistic DeltaK based on the rate of change in the log probability of data between successive K values, we found that STRUCTURE accurately detects the uppermost hierarchical level of structure for the scenarios we tested. As might be expected, the results are sensitive to the type of genetic marker used (AFLP vs. microsatellite), the number of loci scored, the number of populations sampled, and the number of individuals typed in each sample.
Key Findings
1
An ad hoc statistic DeltaK, based on the rate of change in log probability between successive K values, accurately detects the uppermost hierarchical level of structure in the tested scenarios.
2
Detection performance also depends on the number of loci scored, the number of populations sampled, and the number of individuals typed per sample.
3
STRUCTURE's ability to detect K is sensitive to genetic marker type (AFLP vs. microsatellite).
4
The commonly used 'log probability of data' from STRUCTURE often does not correctly estimate the true number of clusters K under non-homogeneous dispersal scenarios.
Research Object
The software STRUCTURE applied to samples of individuals simulated under various dispersal scenarios
Research Subject
Accuracy of detecting the true number of genetic clusters (K), including performance of log probability of data and the DeltaK statistic, across marker types, numbers of loci, populations sampled, and sample sizes
Publication Details
Publication Date
2005-05-11
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai2
Cited by3
Analysis of genetic relationships and evaluation of fruit quality traits in breeding forms of Persian walnut (Juglans regia L.)2026
Allele-defined genome of the autopolyploid sugarcane Saccharum spontaneum L.2018
CLUMPP: a cluster matching and permutation program for dealing with label switching and multimodality in analysis of population structure2007