Error analysis for the improvement of subject ellipsis detection

This paper presents an analysis of the errors of a machine learning method that allow us to propose changes to improve it in future developments. The evaluated system detects Spanish subject ellipsis and yields an accuracy of 85.3%. We extract the wrongly classified instances of our training data (1...

Full description

Bibliographic Details
Authors: Rello, Luz, 1984-, Ferraro, Gabriela, Burga Díaz, Alicia
Format: article
Status:Published version
Publication Date:2011
Country:España
Institution:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
Repository:Recercat. Dipósit de la Recerca de Catalunya
OAI Identifier:oai:recercat.cat:10230/43843
Online Access:http://hdl.handle.net/10230/43843
Access Level:Open access
Keyword:Subject ellipsis
Impersonal construction
Zero pronoun
Error analysis
Linguistic analysis
Machine learning
Elipsis de sujeto
Construcción impersonal
Pronombre zero
Análisis de errores
Análisis lingüístico
Aprendizaje automático
id ES_e0e13ea76b2bf1d1f5a0687278e7de7e
oai_identifier_str oai:recercat.cat:10230/43843
network_acronym_str ES
network_name_str España
repository_id_str
spelling Error analysis for the improvement of subject ellipsis detectionRello, Luz, 1984-Ferraro, GabrielaBurga Díaz, AliciaSubject ellipsisImpersonal constructionZero pronounError analysisLinguistic analysisMachine learningElipsis de sujetoConstrucción impersonalPronombre zeroAnálisis de erroresAnálisis lingüísticoAprendizaje automáticoThis paper presents an analysis of the errors of a machine learning method that allow us to propose changes to improve it in future developments. The evaluated system detects Spanish subject ellipsis and yields an accuracy of 85.3%. We extract the wrongly classified instances of our training data (1,001) and classify the errors. We perform an analysis of these instances taking into account the features and the linguistic patterns involved, which motivate the inclusion of new features and rules in the system.En este trabajo se presenta el análisis de los errores de un método de detección de elipsis de sujeto en español, con el fin de mejorar el sistema en el futuro. El sistema que se evalúa utiliza aprendizaje automático y alcanza una exactitud del 85,3%. El análisis se ha realizado extrayendo de los datos de aprendizaje las instancias que el sistema clasifica erróneamente (1.001), con objeto de establecer una tipología de errores. Cada tipo de error se ha considerado teniendo en cuenta tanto los valores de las características de las instancias como los patrones lingüísticos involucrados. Finalmente, se proponen nuevas características y un conjunto de reglas que puedan aportar una mayor precisión al método.Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN)202020202011info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/43843reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésProcesamiento del lenguaje natural. 2011;(47):223-30© Sociedad Española para el Procesamiento de Lenguaje Natural https://creativecommons.org/licenses/by-nc-nd/4.0/https://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccessoai:recercat.cat:10230/438432026-05-29T05:05:01Z
dc.title.none.fl_str_mv Error analysis for the improvement of subject ellipsis detection
title Error analysis for the improvement of subject ellipsis detection
spellingShingle Error analysis for the improvement of subject ellipsis detection
Rello, Luz, 1984-
Subject ellipsis
Impersonal construction
Zero pronoun
Error analysis
Linguistic analysis
Machine learning
Elipsis de sujeto
Construcción impersonal
Pronombre zero
Análisis de errores
Análisis lingüístico
Aprendizaje automático
title_short Error analysis for the improvement of subject ellipsis detection
title_full Error analysis for the improvement of subject ellipsis detection
title_fullStr Error analysis for the improvement of subject ellipsis detection
title_full_unstemmed Error analysis for the improvement of subject ellipsis detection
title_sort Error analysis for the improvement of subject ellipsis detection
dc.creator.none.fl_str_mv Rello, Luz, 1984-
Ferraro, Gabriela
Burga Díaz, Alicia
author Rello, Luz, 1984-
author_facet Rello, Luz, 1984-
Ferraro, Gabriela
Burga Díaz, Alicia
author_role author
author2 Ferraro, Gabriela
Burga Díaz, Alicia
author2_role author
author
dc.subject.none.fl_str_mv Subject ellipsis
Impersonal construction
Zero pronoun
Error analysis
Linguistic analysis
Machine learning
Elipsis de sujeto
Construcción impersonal
Pronombre zero
Análisis de errores
Análisis lingüístico
Aprendizaje automático
topic Subject ellipsis
Impersonal construction
Zero pronoun
Error analysis
Linguistic analysis
Machine learning
Elipsis de sujeto
Construcción impersonal
Pronombre zero
Análisis de errores
Análisis lingüístico
Aprendizaje automático
description This paper presents an analysis of the errors of a machine learning method that allow us to propose changes to improve it in future developments. The evaluated system detects Spanish subject ellipsis and yields an accuracy of 85.3%. We extract the wrongly classified instances of our training data (1,001) and classify the errors. We perform an analysis of these instances taking into account the features and the linguistic patterns involved, which motivate the inclusion of new features and rules in the system.
publishDate 2011
dc.date.none.fl_str_mv 2011
2020
2020
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10230/43843
url http://hdl.handle.net/10230/43843
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Procesamiento del lenguaje natural. 2011;(47):223-30
dc.rights.none.fl_str_mv https://creativecommons.org/licenses/by-nc-nd/4.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv https://creativecommons.org/licenses/by-nc-nd/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN)
publisher.none.fl_str_mv Sociedad Española para el Procesamiento del Lenguaje Natural (SEPLN)
dc.source.none.fl_str_mv reponame:Recercat. Dipósit de la Recerca de Catalunya
instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
instname_str Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
reponame_str Recercat. Dipósit de la Recerca de Catalunya
collection Recercat. Dipósit de la Recerca de Catalunya
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869422242356527104
score 15,228081