Second-generation PLINK: rising to the challenge of larger and richer datasets
PLINK второго поколения: ответ на вызов, связанный с увеличением объёма и богатства наборов данных
2015-02-24
SCID: 54.1/emjd3gu9
Discuss with AI
PLINK 1.9bit-level parallelismgenome-wide association studiesgenotype likelihoodspopulation genetics
Figures from the paper
Abstract (AI)
BACKGROUND: PLINK 1 is a widely used open-source C/C++ toolset for genome-wide association studies (GWAS) and research in population genetics. However, the steady accumulation of data from imputation and whole-genome sequencing studies has exposed a strong need for faster and scalable implementations of key functions, such as logistic regression, linkage disequilibrium estimation, and genomic distance evaluation. In addition, GWAS and population-genetic data now frequently contain genotype likelihoods, phase information, and/or multiallelic variants, none of which can be represented by PLINK 1's primary data format. FINDINGS: To address these issues, we are developing a second-generation codebase for PLINK. The first major release from this codebase, PLINK 1.9, introduces extensive use of bit-level parallelism, [Formula: see text]-time/constant-space Hardy-Weinberg equilibrium and Fisher's exact tests, and many other algorithmic improvements. In combination, these changes accelerate most operations by 1-4 orders of magnitude, and allow the program to handle datasets too large to fit in RAM. We have also developed an extension to the data format which adds low-overhead support for genotype likelihoods, phase, multiallelic variants, and reference vs. alternate alleles, which is the basis of our planned second release (PLINK 2.0). CONCLUSIONS: The second-generation versions of PLINK will offer dramatic improvements in performance and compatibility. For the first time, users without access to high-end computing resources can perform several essential analyses of the feature-rich and very large genetic datasets coming into use.
Key Findings
1
An extended data format supports genotype likelihoods, phase information, multiallelic variants, and reference-versus-alternate alleles with low overhead.
2
PLINK 1.9 can process datasets too large to fit in RAM, expanding analyses beyond high-end computing environments.
3
PLINK 1.9 introduces bit-level parallelism and algorithmic improvements for scalable genome-wide association and population-genetic analyses.
4
The new implementation accelerates most operations by 1–4 orders of magnitude compared with PLINK 1.
5
The planned PLINK 2.0 release is based on this richer format and aims to improve compatibility with increasingly complex genetic datasets.
Research Object
Second-generation PLINK software and its genetic datasets for genome-wide association and population-genetic analyses
Research Subject
Scalable computational performance and data-format compatibility for analyzing very large, feature-rich genetic datasets, including genotype likelihoods, phase information, and multiallelic variants
Publication Details
Publication Date
2015-02-24
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai1
Cited by3
Analysis of 173,303 exomes and genomes in the Pakistan Genome Resource2026
Multi-Genome-Wide association studies provide new insights into the genetic architecture of Varroa resistance in honeybees2025
Introgression of Eastern Chinese and Southern Chinese haplotypes contributes to the improvement of fertility and immunity in European modern pigs2020