Genome annotation assessment in Drosophila melanogaster.

Computational methods for automated genome annotation are critical to our community's ability to make full use of the large volume of genomic sequence being generated and released. To explore the accuracy of these automated feature prediction tools in the genomes of higher organisms, we evaluat...

ver descrição completa

Detalhes bibliográficos
Autores: Reese, Martin G., Hartzell, George, Harris, Nomi L., Ohler, Uwe, Abril Ferrando, Josep Francesc, 1970-, Lewis, Suzanna E.
Formato: artículo
Estado:Versión publicada
Fecha de publicación:2000
País:España
Recursos:Universidad de Barcelona
Repositorio:Dipòsit Digital de la UB
OAI Identifier:oai:diposit.ub.edu:2445/192625
Acesso em linha:https://hdl.handle.net/2445/192625
Access Level:acceso abierto
Palavra-chave:Drosòfila melanogaster
Drosòfila
Genòmica
Drosophila melanogaster
Drosophila
Genomics
id ES_e9c1cbc96ee0f93ae89bccb6f0435fac
oai_identifier_str oai:diposit.ub.edu:2445/192625
network_acronym_str ES
network_name_str España
repository_id_str
spelling Genome annotation assessment in Drosophila melanogaster.Reese, Martin G.Hartzell, GeorgeHarris, Nomi L.Ohler, UweAbril Ferrando, Josep Francesc, 1970-Lewis, Suzanna E.Drosòfila melanogasterDrosòfilaGenòmicaDrosophila melanogasterDrosophilaGenomicsComputational methods for automated genome annotation are critical to our community's ability to make full use of the large volume of genomic sequence being generated and released. To explore the accuracy of these automated feature prediction tools in the genomes of higher organisms, we evaluated their performance on a large, well-characterized sequence contig from the Adh region ofDrosophila melanogaster. This experiment, known as the Genome Annotation Assessment Project (GASP), was launched in May 1999. Twelve groups, applying state-of-the-art tools, contributed predictions for features including gene structure, protein homologies, promoter sites, and repeat elements. We evaluated these predictions using two standards, one based on previously unreleased high-quality full-length cDNA sequences and a second based on the set of annotations generated as part of an in-depth study of the region by a group ofDrosophila experts. Although these standard sets only approximate the unknown distribution of features in this region, we believe that when taken in context the results of an evaluation based on them are meaningful. The results were presented as a tutorial at the conference on Intelligent Systems in Molecular Biology (ISMB-99) in August 1999. Over 95% of the coding nucleotides in the region were correctly identified by the majority of the gene finders, and the correct intron/exon structures were predicted for >40% of the genes. Homology-based annotation techniques recognized and associated functions with almost half of the genes in the region; the remainder were only identified by the ab initio techniques. This experiment also presents the first assessment of promoter prediction techniques for a significant number of genes in a large contiguous region. We discovered that the promoter predictors' high false-positive rates make their predictions difficult to use. Integrating gene finding and cDNA/EST alignments with promoter predictions decreases the number of false-positive classifications but discovers less than one-third of the promoters in the region. We believe that by establishing standards for evaluating genomic annotations and by assessing the performance of existing automated genome annotation tools, this experiment establishes a baseline that contributes to the value of ongoing large-scale annotation projects and should guide further research in genome informatics.Cold Spring Harbor Laboratory Press2000info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfhttps://hdl.handle.net/2445/192625Articles publicats en revistes (Genètica, Microbiologia i Estadística)reponame:Dipòsit Digital de la UBinstname:Universidad de BarcelonaInglésReproducció del document publicat a: https://doi.org/10.1101/gr.10.4.483Genome Research, 2000, vol. 10, num. 4, p. 483-501https://doi.org/10.1101/gr.10.4.483cc-by-nc (c) Reese, Martin G. et al., 2000https://creativecommons.org/licenses/by-nc/4.0/info:eu-repo/semantics/openAccessoai:diposit.ub.edu:2445/1926252026-05-27T06:46:51Z
dc.title.none.fl_str_mv Genome annotation assessment in Drosophila melanogaster.
title Genome annotation assessment in Drosophila melanogaster.
spellingShingle Genome annotation assessment in Drosophila melanogaster.
Reese, Martin G.
Drosòfila melanogaster
Drosòfila
Genòmica
Drosophila melanogaster
Drosophila
Genomics
title_short Genome annotation assessment in Drosophila melanogaster.
title_full Genome annotation assessment in Drosophila melanogaster.
title_fullStr Genome annotation assessment in Drosophila melanogaster.
title_full_unstemmed Genome annotation assessment in Drosophila melanogaster.
title_sort Genome annotation assessment in Drosophila melanogaster.
dc.creator.none.fl_str_mv Reese, Martin G.
Hartzell, George
Harris, Nomi L.
Ohler, Uwe
Abril Ferrando, Josep Francesc, 1970-
Lewis, Suzanna E.
author Reese, Martin G.
author_facet Reese, Martin G.
Hartzell, George
Harris, Nomi L.
Ohler, Uwe
Abril Ferrando, Josep Francesc, 1970-
Lewis, Suzanna E.
author_role author
author2 Hartzell, George
Harris, Nomi L.
Ohler, Uwe
Abril Ferrando, Josep Francesc, 1970-
Lewis, Suzanna E.
author2_role author
author
author
author
author
dc.subject.none.fl_str_mv Drosòfila melanogaster
Drosòfila
Genòmica
Drosophila melanogaster
Drosophila
Genomics
topic Drosòfila melanogaster
Drosòfila
Genòmica
Drosophila melanogaster
Drosophila
Genomics
description Computational methods for automated genome annotation are critical to our community's ability to make full use of the large volume of genomic sequence being generated and released. To explore the accuracy of these automated feature prediction tools in the genomes of higher organisms, we evaluated their performance on a large, well-characterized sequence contig from the Adh region ofDrosophila melanogaster. This experiment, known as the Genome Annotation Assessment Project (GASP), was launched in May 1999. Twelve groups, applying state-of-the-art tools, contributed predictions for features including gene structure, protein homologies, promoter sites, and repeat elements. We evaluated these predictions using two standards, one based on previously unreleased high-quality full-length cDNA sequences and a second based on the set of annotations generated as part of an in-depth study of the region by a group ofDrosophila experts. Although these standard sets only approximate the unknown distribution of features in this region, we believe that when taken in context the results of an evaluation based on them are meaningful. The results were presented as a tutorial at the conference on Intelligent Systems in Molecular Biology (ISMB-99) in August 1999. Over 95% of the coding nucleotides in the region were correctly identified by the majority of the gene finders, and the correct intron/exon structures were predicted for >40% of the genes. Homology-based annotation techniques recognized and associated functions with almost half of the genes in the region; the remainder were only identified by the ab initio techniques. This experiment also presents the first assessment of promoter prediction techniques for a significant number of genes in a large contiguous region. We discovered that the promoter predictors' high false-positive rates make their predictions difficult to use. Integrating gene finding and cDNA/EST alignments with promoter predictions decreases the number of false-positive classifications but discovers less than one-third of the promoters in the region. We believe that by establishing standards for evaluating genomic annotations and by assessing the performance of existing automated genome annotation tools, this experiment establishes a baseline that contributes to the value of ongoing large-scale annotation projects and should guide further research in genome informatics.
publishDate 2000
dc.date.none.fl_str_mv 2000
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv https://hdl.handle.net/2445/192625
url https://hdl.handle.net/2445/192625
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Reproducció del document publicat a: https://doi.org/10.1101/gr.10.4.483
Genome Research, 2000, vol. 10, num. 4, p. 483-501
https://doi.org/10.1101/gr.10.4.483
dc.rights.none.fl_str_mv cc-by-nc (c) Reese, Martin G. et al., 2000
https://creativecommons.org/licenses/by-nc/4.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv cc-by-nc (c) Reese, Martin G. et al., 2000
https://creativecommons.org/licenses/by-nc/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Cold Spring Harbor Laboratory Press
publisher.none.fl_str_mv Cold Spring Harbor Laboratory Press
dc.source.none.fl_str_mv Articles publicats en revistes (Genètica, Microbiologia i Estadística)
reponame:Dipòsit Digital de la UB
instname:Universidad de Barcelona
instname_str Universidad de Barcelona
reponame_str Dipòsit Digital de la UB
collection Dipòsit Digital de la UB
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869423079436845056
score 15,301629