The SEED and the Rapid Annotation of microbial genomes using Subsystems Technology (RAST)

SEED и быстрая аннотация микробных геномов с использованием технологии подсистем (RAST)
Ross Overbeek, Robert Olson, Gordon D. Pusch, G J Olsen, James J. Davis, Terry Disz, Robert A. Edwards, Svetlana Gerdes, Bruce Parrello, Maulik Shukla, Veronika Vonstein, Alice R. Wattam, Fangfang Xia, Rick Stevens
2013-11-29

FIGfam protein familiesRAST annotation pipelineSEED databaseSubsystems Technologymicrobial genome annotation
In 2004, the SEED (http://pubseed.theseed.org/) was created to provide consistent and accurate genome annotations across thousands of genomes and as a platform for discovering and developing de novo annotations. The SEED is a constantly updated integration of genomic data with a genome database, web front end, API and server scripts. It is used by many scientists for predicting gene functions and discovering new pathways. In addition to being a powerful database for bioinformatics research, the SEED also houses subsystems (collections of functionally related protein families) and their derived FIGfams (protein families), which represent the core of the RAST annotation engine (http://rast.nmpdr.org/). When a new genome is submitted to RAST, genes are called and their annotations are made by comparison to the FIGfam collection. If the genome is made public, it is then housed within the SEED and its proteins populate the FIGfam collection. This annotation cycle has proven to be a robust and scalable solution to the problem of annotating the exponentially increasing number of genomes. To date, >12 000 users worldwide have annotated >60 000 distinct genomes using RAST. Here we describe the interconnectedness of the SEED database and RAST, the RAST annotation pipeline and updates to both resources.
1
Publicly released RAST genomes are incorporated into SEED, and their proteins expand the FIGfam collection, creating a continuous annotation-improvement cycle.
2
RAST annotates submitted genomes by calling genes and comparing their predicted proteins with the SEED-derived FIGfam protein-family collection.
3
SEED integrates genomic data, a genome database, web interface, API, server scripts, subsystems, and FIGfams for genome annotation and pathway discovery.
4
The SEED was created in 2004 to provide consistent, accurate genome annotations and support discovery of de novo annotations across thousands of genomes.
5
The SEED–RAST system has provided a robust, scalable approach: more than 12,000 users have annotated over 60,000 distinct genomes worldwide.

The SEED genomic database and RAST microbial genome annotation system

The interconnected genome-annotation workflow, scalability, and ongoing updates linking SEED data, subsystems/FIGfams, and the RAST annotation pipeline

Publication Details
Publication Date
2013-11-29
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Ross Overbeek
Robert Olson
Gordon D. Pusch
G J Olsen
James J. Davis
Terry Disz
Robert A. Edwards
Svetlana Gerdes
Bruce Parrello
Maulik Shukla
Veronika Vonstein
Alice R. Wattam
Fangfang Xia
Rick Stevens
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%