Improving genome-wide scans of positive selection by using protein isoforms of similar length

Large-scale evolutionary studies often require the automated construction of alignments of a large number of homologous gene families. The majority of eukaryotic genes can produce different transcripts due to alternative splicing or transcription initiation, and many such transcripts encode differen...

Descripción completa

Detalles Bibliográficos
Autores: Villanueva Cañas, José Luis, 1984-, Laurie, Steven, 1973-, Albà Soler, Mar
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2013
País:España
Institución:Universitat Pompeu Fabra
Repositorio:Repositorio Digital de la UPF
OAI Identifier:oai:repositori.upf.edu:10230/27682
Acceso en línea:http://hdl.handle.net/10230/27682
http://dx.doi.org/10.1093/gbe/evt017
Access Level:acceso abierto
Palabra clave:Evolució molecular
Selecció natural
Genètica
Proteïnes
id ES_e1084c58e4fcefcc4edbc694e5a4acbf
oai_identifier_str oai:repositori.upf.edu:10230/27682
network_acronym_str ES
network_name_str España
repository_id_str
spelling Improving genome-wide scans of positive selection by using protein isoforms of similar lengthVillanueva Cañas, José Luis, 1984-Laurie, Steven, 1973-Albà Soler, MarEvolució molecularSelecció naturalGenèticaProteïnesLarge-scale evolutionary studies often require the automated construction of alignments of a large number of homologous gene families. The majority of eukaryotic genes can produce different transcripts due to alternative splicing or transcription initiation, and many such transcripts encode different protein isoforms. As analyses tend to be gene centered, one single-protein isoform per gene is selected for the alignment, with the de facto approach being to use the longest protein isoform per gene (Longest), presumably to avoid including partial sequences and to maximize sequence information. Here, we show that this approach is problematic because it increases the number of indels in the alignments due to the inclusion of nonhomologous regions, such as those derived from species-specific exons, increasing the number of misaligned positions. With the aim of ameliorating this problem, we have developed a novel heuristic, Protein ALignment Optimizer (PALO), which, for each gene family, selects the combination of protein isoforms that are most similar in length. We examine several evolutionary parameters inferred from alignments in which the only difference is the method used to select the protein isoform combination: Longest, PALO, the combination that results in the highest sequence conservation, and a randomly selected combination. We observe that Longest tends to overestimate both nonsynonymous and synonymous substitution rates when compared with PALO, which is most likely due to an excess of misaligned positions. The estimation of the fraction of genes that have experienced positive selection by maximum likelihood is very sensitive to the method of isoform selection employed, both when alignments are constructed with MAFFT and with Prank(+F). Longest performs better than a random combination but still estimates up to 3 times more positively selected genes than the combination showing the highest conservation, indicating the presence of many false positives. We show that PALO can eliminate the majority of such false positives and thus that it is a more appropriate approach for large-scale analyses than Longest. A web server has been set up to facilitate the use of PALO given a user-defined set of gene families; it is available at http://evolutionarygenomics.imim.es/palo.This work was funded by Ministerio de Economía y Competitividad (FPI BES-2010-038494 to J.L.V.-C., Plan Nacional BIO2009-08160 and BFU2012-36820) and Fundació ICREA to M.M.A.Oxford University Press201620162013info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/27682http://dx.doi.org/10.1093/gbe/evt017reponame:Repositorio Digital de la UPFinstname:Universitat Pompeu FabraInglésGenome Biology and Evolution. 2013;5(2):457-67info:eu-repo/grantAgreement/ES/3PN/BES2010-038494info:eu-repo/grantAgreement/ES/3PN/BIO2009-08160info:eu-repo/grantAgreement/ES/3PN/BFU2012-36820© José Luis Villanueva-Cañas, Steve Laurie and M. Mar Albà. 2013. Published by Oxford University Press. This is an Open Access article distributed under the terms of a Creative Commons Attribution Licensehttp://creativecommons.org/licenses/by-nc/3.0/info:eu-repo/semantics/openAccessoai:repositori.upf.edu:10230/276822026-06-12T07:21:37Z
dc.title.none.fl_str_mv Improving genome-wide scans of positive selection by using protein isoforms of similar length
title Improving genome-wide scans of positive selection by using protein isoforms of similar length
spellingShingle Improving genome-wide scans of positive selection by using protein isoforms of similar length
Villanueva Cañas, José Luis, 1984-
Evolució molecular
Selecció natural
Genètica
Proteïnes
title_short Improving genome-wide scans of positive selection by using protein isoforms of similar length
title_full Improving genome-wide scans of positive selection by using protein isoforms of similar length
title_fullStr Improving genome-wide scans of positive selection by using protein isoforms of similar length
title_full_unstemmed Improving genome-wide scans of positive selection by using protein isoforms of similar length
title_sort Improving genome-wide scans of positive selection by using protein isoforms of similar length
dc.creator.none.fl_str_mv Villanueva Cañas, José Luis, 1984-
Laurie, Steven, 1973-
Albà Soler, Mar
author Villanueva Cañas, José Luis, 1984-
author_facet Villanueva Cañas, José Luis, 1984-
Laurie, Steven, 1973-
Albà Soler, Mar
author_role author
author2 Laurie, Steven, 1973-
Albà Soler, Mar
author2_role author
author
dc.subject.none.fl_str_mv Evolució molecular
Selecció natural
Genètica
Proteïnes
topic Evolució molecular
Selecció natural
Genètica
Proteïnes
description Large-scale evolutionary studies often require the automated construction of alignments of a large number of homologous gene families. The majority of eukaryotic genes can produce different transcripts due to alternative splicing or transcription initiation, and many such transcripts encode different protein isoforms. As analyses tend to be gene centered, one single-protein isoform per gene is selected for the alignment, with the de facto approach being to use the longest protein isoform per gene (Longest), presumably to avoid including partial sequences and to maximize sequence information. Here, we show that this approach is problematic because it increases the number of indels in the alignments due to the inclusion of nonhomologous regions, such as those derived from species-specific exons, increasing the number of misaligned positions. With the aim of ameliorating this problem, we have developed a novel heuristic, Protein ALignment Optimizer (PALO), which, for each gene family, selects the combination of protein isoforms that are most similar in length. We examine several evolutionary parameters inferred from alignments in which the only difference is the method used to select the protein isoform combination: Longest, PALO, the combination that results in the highest sequence conservation, and a randomly selected combination. We observe that Longest tends to overestimate both nonsynonymous and synonymous substitution rates when compared with PALO, which is most likely due to an excess of misaligned positions. The estimation of the fraction of genes that have experienced positive selection by maximum likelihood is very sensitive to the method of isoform selection employed, both when alignments are constructed with MAFFT and with Prank(+F). Longest performs better than a random combination but still estimates up to 3 times more positively selected genes than the combination showing the highest conservation, indicating the presence of many false positives. We show that PALO can eliminate the majority of such false positives and thus that it is a more appropriate approach for large-scale analyses than Longest. A web server has been set up to facilitate the use of PALO given a user-defined set of gene families; it is available at http://evolutionarygenomics.imim.es/palo.
publishDate 2013
dc.date.none.fl_str_mv 2013
2016
2016
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10230/27682
http://dx.doi.org/10.1093/gbe/evt017
url http://hdl.handle.net/10230/27682
http://dx.doi.org/10.1093/gbe/evt017
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Genome Biology and Evolution. 2013;5(2):457-67
info:eu-repo/grantAgreement/ES/3PN/BES2010-038494
info:eu-repo/grantAgreement/ES/3PN/BIO2009-08160
info:eu-repo/grantAgreement/ES/3PN/BFU2012-36820
dc.rights.none.fl_str_mv http://creativecommons.org/licenses/by-nc/3.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv http://creativecommons.org/licenses/by-nc/3.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Oxford University Press
publisher.none.fl_str_mv Oxford University Press
dc.source.none.fl_str_mv reponame:Repositorio Digital de la UPF
instname:Universitat Pompeu Fabra
instname_str Universitat Pompeu Fabra
reponame_str Repositorio Digital de la UPF
collection Repositorio Digital de la UPF
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869422256051978240
score 15,228081