Search and clustering orders of magnitude faster than BLAST
Поиск и кластеризация с производительностью на несколько порядков выше, чем у BLAST
2010-08-12
SCID: 54.1/kkq4k7nk
Discuss with AI
UBLASTUCLUSTUSEARCHfast sequence searchhigh-throughput sequence clustering
Figures from the paper
Abstract (AI)
MOTIVATION: Biological sequence data is accumulating rapidly, motivating the development of improved high-throughput methods for sequence classification. RESULTS: UBLAST and USEARCH are new algorithms enabling sensitive local and global search of large sequence databases at exceptionally high speeds. They are often orders of magnitude faster than BLAST in practical applications, though sensitivity to distant protein relationships is lower. UCLUST is a new clustering method that exploits USEARCH to assign sequences to clusters. UCLUST offers several advantages over the widely used program CD-HIT, including higher speed, lower memory use, improved sensitivity, clustering at lower identities and classification of much larger datasets. AVAILABILITY: Binaries are available at no charge for non-commercial use at http://www.drive5.com/usearch.
Key Findings
1
UBLAST and USEARCH are often orders of magnitude faster than BLAST in practical applications.
2
UBLAST and USEARCH have lower sensitivity to distant protein relationships compared to BLAST.
3
UBLAST and USEARCH provide sensitive local and global sequence search of large databases at exceptionally high speeds.
4
UCLUST can classify much larger datasets than CD-HIT.
5
UCLUST uses USEARCH to assign sequences to clusters and offers higher speed than CD-HIT.
6
UCLUST uses less memory than CD-HIT, provides improved sensitivity, and supports clustering at lower identity thresholds.
Research Object
Algorithms UBLAST, USEARCH and UCLUST for biological sequence search and clustering
Research Subject
Performance and trade-offs of these algorithms: search and clustering speed (orders-of-magnitude faster than BLAST), sensitivity (including reduced sensitivity for distant protein relationships), memory usage, ability to cluster at lower identity thresholds, and scalability to large datasets
Publication Details
Publication Date
2010-08-12
Journal
Publisher
ISSN
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest