Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences
Cd-hit: быстрая программа для кластеризации и сравнения больших наборов белковых или нуклеотидных последовательностей
2006-05-26
SCID: 54.1/9vd25ef8
Discuss with AI
BLASTCD-HITcd-hit-2dcd-hit-estcd-hit-est-2dlarge-scale sequence databasesnucleotide sequence clusteringprotein sequence clusteringsequence comparisonultrafast clustering
Figures from the paper
Abstract (AI)
MOTIVATION: In 2001 and 2002, we published two papers (Bioinformatics, 17, 282-283, Bioinformatics, 18, 77-82) describing an ultrafast protein sequence clustering program called cd-hit. This program can efficiently cluster a huge protein database with millions of sequences. However, the applications of the underlying algorithm are not limited to only protein sequences clustering, here we present several new programs using the same algorithm including cd-hit-2d, cd-hit-est and cd-hit-est-2d. Cd-hit-2d compares two protein datasets and reports similar matches between them; cd-hit-est clusters a DNA/RNA sequence database and cd-hit-est-2d compares two nucleotide datasets. All these programs can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on the popular sequence comparison and database search tools, such as BLAST.
Key Findings
1
All cd-hit variants can handle huge datasets with millions of sequences and can be hundreds of times faster than methods based on tools like BLAST.
2
The cd-hit algorithm is generalized into new programs: cd-hit-2d for comparing two protein datasets and reporting similar matches.
3
cd-hit is an ultrafast program that can efficiently cluster very large protein databases containing millions of sequences.
4
cd-hit-est extends the algorithm to cluster DNA/RNA (nucleotide) sequence databases.
5
cd-hit-est-2d compares two nucleotide datasets to report similar matches between them.
Research Object
CD-HIT suite of programs for clustering and comparing large sets of protein or nucleotide sequences
Research Subject
Efficient clustering and pairwise comparison of very large protein and nucleotide sequence datasets (scalability, speed, and similarity-match reporting compared to standard tools)
Publication Details
Publication Date
2006-05-26
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by3
A viral ORFeome library for systems-level genetic dissection of host-pathogen interactions2026
Discovery of specific rhizosphere bacteria <i>Rhodanobacter</i> involved in KAI2-mediated drought tolerance in <i>Arabidopsis</i>2026
PathogenFinder - Distinguishing Friend from Foe Using Bacterial Whole Genome Sequence Data2013