Text mining and expert curation to develop a database on psychiatric diseases and their genes

Vincent Warnault, Ferrán Sanz, Jesús Giraldo, Olga Valverde, Laura I. Furlong, Álex Bravo, Alba Gutiérrez‐Sacristán, Marta Portero Tresserra, Antonio Armario, M. Carmen Blanco-Gandía, Adriana Farré, Lierni Fernández‐Ibarrondo, Francina Fonseca, Ángela Leis, Anna Mané, Miguel Ángel Mayer, Sandra Montagud‐Romero, Roser Nadal, Jordí Ortiz, Francisco Javier Pavón, Ezequiel Jesús Perez, Marta Rodríguez‐Arias, Antonia Serrano, Marta Torrens
2017-06-26

SCID:  54.1/zfdw5uka
Psychiatric disorders constitute one of the main causes of disability worldwide. During the past years, considerable research has been conducted on the genetic architecture of such diseases, although little understanding of their etiology has been achieved. The difficulty to access up-to-date, relevant genotype-phenotype information has hampered the application of this wealth of knowledge to translational research and clinical practice in order to improve diagnosis and treatment of psychiatric patients. PsyGeNET (http://www.psygenet.org/) has been developed with the aim of supporting research on the genetic architecture of psychiatric diseases, by providing integrated and structured accessibility to their genotype–phenotype association data, together with analysis and visualization tools. In this article, we describe the protocol developed for the sustainable update of this knowledge resource. It includes the recruitment of a team of domain experts in order to perform the curation of the data extracted by text mining. Annotation guidelines and a web-based annotation tool were developed to support the curators’ tasks. A curation workflow was designed including a pilot phase and two rounds of curation and analysis phases. Negative evidence from the literature on gene–disease associations (GDAs) was taken into account in the curation process. We report the results of the application of this workflow to the curation of GDAs for PsyGeNET, including the analysis of the inter-annotator agreement and suggest this model as a suitable approach for the sustainable development and update of knowledge resources. Database URL: http://www.psygenet.org PsyGeNET corpus: http://www.psygenet.org/ds/PsyGeNET/results/psygenetCorpus.tar
Publication Details
Publication Date
2017-06-26
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Vincent Warnault
Ferrán Sanz
Jesús Giraldo
Olga Valverde
Laura I. Furlong
Álex Bravo
Alba Gutiérrez‐Sacristán
Marta Portero Tresserra
Antonio Armario
M. Carmen Blanco-Gandía
Adriana Farré
Lierni Fernández‐Ibarrondo
Francina Fonseca
Ángela Leis
Anna Mané
Miguel Ángel Mayer
Sandra Montagud‐Romero
Roser Nadal
Jordí Ortiz
Francisco Javier Pavón
Ezequiel Jesús Perez
Marta Rodríguez‐Arias
Antonia Serrano
Marta Torrens
Explore More Research
Use the citation graph to discover related papers and expand your research horizons.
Click any node to explore
Download PDF
100%