Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences

Cd-hit: быстрая программа для кластеризации и сравнения больших наборов белковых или нуклеотидных последовательностей
Weizhong Li, Adam Godzik
2006-05-26

BLASTCD-HITcd-hit-2dcd-hit-estcd-hit-est-2dlarge-scale sequence databasesnucleotide sequence clusteringprotein sequence clusteringsequence comparisonultrafast clustering
MOTIVATION: In 2001 and 2002, we published two papers (Bioinformatics, 17, 282-283, Bioinformatics, 18, 77-82) describing an ultrafast protein sequence clustering program called cd-hit. This program can efficiently cluster a huge protein database with millions of sequences. However, the applications of the underlying algorithm are not limited to only protein sequences clustering, here we present several new programs using the same algorithm including cd-hit-2d, cd-hit-est and cd-hit-est-2d. Cd-hit-2d compares two protein datasets and reports similar matches between them; cd-hit-est clusters a DNA/RNA sequence database and cd-hit-est-2d compares two nucleotide datasets. All these programs can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on the popular sequence comparison and database search tools, such as BLAST.
1
All cd-hit variants can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on tools like BLAST.
2
The cd-hit algorithm is generalized into new programs: cd-hit-2d for comparing two protein datasets and reporting similar matches.
3
cd-hit is an ultrafast program that can efficiently cluster very large protein databases containing millions of sequences.
4
cd-hit-est extends the algorithm to cluster DNA/RNA (nucleotide) sequence databases.
5
cd-hit-est-2d compares two nucleotide datasets to report similar matches between them.

CD-HIT suite of programs for clustering and comparing large sets of protein or nucleotide sequences

Efficient clustering and pairwise comparison of very large protein and nucleotide sequence datasets (scalability, speed, and similarity-match reporting compared to standard tools)

Publication Details
Publication Date
2006-05-26
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Weizhong Li
Adam Godzik
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%