Detecting the number of clusters of individuals using the software <scp>structure</scp>: a simulation study

Guillaume Evanno, Sébastien Regnaut, Jérôme Goudet
2005-05-11

AFLP vs. microsatelliteDeltaK statisticSTRUCTURE softwareindividual-based simulationnumber of clusters (K)
The identification of genetically homogeneous groups of individuals is a long standing issue in population genetics. A recent Bayesian algorithm implemented in the software STRUCTURE allows the identification of such groups. However, the ability of this algorithm to detect the true number of clusters (K) in a sample of individuals when patterns of dispersal among populations are not homogeneous has not been tested. The goal of this study is to carry out such tests, using various dispersal scenarios from data generated with an individual-based model. We found that in most cases the estimated 'log probability of data' does not provide a correct estimation of the number of clusters, K. However, using an ad hoc statistic DeltaK based on the rate of change in the log probability of data between successive K values, we found that STRUCTURE accurately detects the uppermost hierarchical level of structure for the scenarios we tested. As might be expected, the results are sensitive to the type of genetic marker used (AFLP vs. microsatellite), the number of loci scored, the number of populations sampled, and the number of individuals typed in each sample.
1
An ad hoc statistic DeltaK, based on the rate of change in log probability between successive K values, accurately detects the uppermost hierarchical level of structure in the tested scenarios.
2
Detection performance also depends on the number of loci scored, the number of populations sampled, and the number of individuals typed per sample.
3
STRUCTURE's ability to detect K is sensitive to genetic marker type (AFLP vs. microsatellite).
4
The commonly used 'log probability of data' from STRUCTURE often does not correctly estimate the true number of clusters K under non-homogeneous dispersal scenarios.

The software STRUCTURE applied to samples of individuals simulated under various dispersal scenarios

Accuracy of detecting the true number of genetic clusters (K), including performance of log probability of data and the DeltaK statistic, across marker types, numbers of loci, populations sampled, and sample sizes

Publication Details
Publication Date
2005-05-11
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Guillaume Evanno
Sébastien Regnaut
Jérôme Goudet
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%