Error analysis for the improvement of subject ellipsis detection
This paper presents an analysis of the errors of a machine learning method that allow us to propose changes to improve it in future developments. The evaluated system detects Spanish subject ellipsis and yields an accuracy of 85.3%. We extract the wrongly classified instances of our training data (1...
| Authors: | , , |
|---|---|
| Format: | article |
| Status: | Published version |
| Publication Date: | 2011 |
| Country: | España |
| Institution: | Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| Repository: | Recercat. Dipósit de la Recerca de Catalunya |
| OAI Identifier: | oai:recercat.cat:10230/43843 |
| Online Access: | http://hdl.handle.net/10230/43843 |
| Access Level: | Open access |
| Keyword: | Subject ellipsis Impersonal construction Zero pronoun Error analysis Linguistic analysis Machine learning Elipsis de sujeto Construcción impersonal Pronombre zero Análisis de errores Análisis lingüístico Aprendizaje automático |
| id |
ES_e0e13ea76b2bf1d1f5a0687278e7de7e |
|---|---|
| oai_identifier_str |
oai:recercat.cat:10230/43843 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Error analysis for the improvement of subject ellipsis detectionRello, Luz, 1984-Ferraro, GabrielaBurga Díaz, AliciaSubject ellipsisImpersonal constructionZero pronounError analysisLinguistic analysisMachine learningElipsis de sujetoConstrucción impersonalPronombre zeroAnálisis de erroresAnálisis lingüísticoAprendizaje automáticoThis paper presents an analysis of the errors of a machine learning method that allow us to propose changes to improve it in future developments. The evaluated system detects Spanish subject ellipsis and yields an accuracy of 85.3%. We extract the wrongly classified instances of our training data (1,001) and classify the errors. We perform an analysis of these instances taking into account the features and the linguistic patterns involved, which motivate the inclusion of new features and rules in the system.En este trabajo se presenta el análisis de los errores de un método de detección de elipsis de sujeto en español, con el fin de mejorar el sistema en el futuro. El sistema que se evalúa utiliza aprendizaje automático y alcanza una exactitud del 85,3%. El análisis se ha realizado extrayendo de los datos de aprendizaje las instancias que el sistema clasifica erróneamente (1.001), con objeto de establecer una tipología de errores. Cada tipo de error se ha considerado teniendo en cuenta tanto los valores de las características de las instancias como los patrones lingüísticos involucrados. Finalmente, se proponen nuevas características y un conjunto de reglas que puedan aportar una mayor precisión al método.Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN)202020202011info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/43843reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésProcesamiento del lenguaje natural. 2011;(47):223-30© Sociedad Española para el Procesamiento de Lenguaje Natural https://creativecommons.org/licenses/by-nc-nd/4.0/https://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccessoai:recercat.cat:10230/438432026-05-29T05:05:01Z |
| dc.title.none.fl_str_mv |
Error analysis for the improvement of subject ellipsis detection |
| title |
Error analysis for the improvement of subject ellipsis detection |
| spellingShingle |
Error analysis for the improvement of subject ellipsis detection Rello, Luz, 1984- Subject ellipsis Impersonal construction Zero pronoun Error analysis Linguistic analysis Machine learning Elipsis de sujeto Construcción impersonal Pronombre zero Análisis de errores Análisis lingüístico Aprendizaje automático |
| title_short |
Error analysis for the improvement of subject ellipsis detection |
| title_full |
Error analysis for the improvement of subject ellipsis detection |
| title_fullStr |
Error analysis for the improvement of subject ellipsis detection |
| title_full_unstemmed |
Error analysis for the improvement of subject ellipsis detection |
| title_sort |
Error analysis for the improvement of subject ellipsis detection |
| dc.creator.none.fl_str_mv |
Rello, Luz, 1984- Ferraro, Gabriela Burga Díaz, Alicia |
| author |
Rello, Luz, 1984- |
| author_facet |
Rello, Luz, 1984- Ferraro, Gabriela Burga Díaz, Alicia |
| author_role |
author |
| author2 |
Ferraro, Gabriela Burga Díaz, Alicia |
| author2_role |
author author |
| dc.subject.none.fl_str_mv |
Subject ellipsis Impersonal construction Zero pronoun Error analysis Linguistic analysis Machine learning Elipsis de sujeto Construcción impersonal Pronombre zero Análisis de errores Análisis lingüístico Aprendizaje automático |
| topic |
Subject ellipsis Impersonal construction Zero pronoun Error analysis Linguistic analysis Machine learning Elipsis de sujeto Construcción impersonal Pronombre zero Análisis de errores Análisis lingüístico Aprendizaje automático |
| description |
This paper presents an analysis of the errors of a machine learning method that allow us to propose changes to improve it in future developments. The evaluated system detects Spanish subject ellipsis and yields an accuracy of 85.3%. We extract the wrongly classified instances of our training data (1,001) and classify the errors. We perform an analysis of these instances taking into account the features and the linguistic patterns involved, which motivate the inclusion of new features and rules in the system. |
| publishDate |
2011 |
| dc.date.none.fl_str_mv |
2011 2020 2020 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/publishedVersion |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10230/43843 |
| url |
http://hdl.handle.net/10230/43843 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
Procesamiento del lenguaje natural. 2011;(47):223-30 |
| dc.rights.none.fl_str_mv |
https://creativecommons.org/licenses/by-nc-nd/4.0/ info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
https://creativecommons.org/licenses/by-nc-nd/4.0/ |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf application/pdf |
| dc.publisher.none.fl_str_mv |
Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN) |
| publisher.none.fl_str_mv |
Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN) |
| dc.source.none.fl_str_mv |
reponame:Recercat. Dipósit de la Recerca de Catalunya instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| instname_str |
Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| reponame_str |
Recercat. Dipósit de la Recerca de Catalunya |
| collection |
Recercat. Dipósit de la Recerca de Catalunya |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869422242356527104 |
| score |
15,228081 |