Feature analysis for discriminative confidence estimation in spoken term detection

This is the author’s version of a work that was accepted for publication in Computer Speech & Language. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Cha...

Descripción completa

Detalles Bibliográficos
Autores: Tejedor Noguerales, Javier, Toledano, Doroteo T., Wang, Dong, King, Simon, Colás Pasamontes, José
Tipo de recurso: artículo
Fecha de publicación:2014
País:España
Institución:Universidad Autónoma de Madrid
Repositorio:Biblos-e Archivo. Repositorio Institucional de la UAM
Idioma:inglés
OAI Identifier:oai:repositorio.uam.es:10486/662781
Acceso en línea:http://hdl.handle.net/10486/662781
https://dx.doi.org/10.1016/j.csl.2013.09.008
Access Level:acceso abierto
Palabra clave:Discriminative confidence
Feature analysis
Speech recognition
Spoken term detection
Telecomunicaciones
id ES_a448de7a2b255785d5b186a80002fbf6
oai_identifier_str oai:repositorio.uam.es:10486/662781
network_acronym_str ES
network_name_str España
repository_id_str
spelling Feature analysis for discriminative confidence estimation in spoken term detectionTejedor Noguerales, JavierToledano, Doroteo T.Wang, DongKing, SimonColás Pasamontes, JoséDiscriminative confidenceFeature analysisSpeech recognitionSpoken term detectionTelecomunicacionesThis is the author’s version of a work that was accepted for publication in Computer Speech & Language. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in Computer Speech & Language, 28, 5, (2014) DOI: 10.1016/j.csl.2013.09.008Discriminative confidence based on multi-layer perceptrons (MLPs) and multiple features has shown significant advantage compared to the widely used lattice-based confidence in spoken term detection (STD). Although the MLP-based framework can handle any features derived from a multitude of sources, choosing all possible features may lead to over complex models and hence less generality. In this paper, we design an extensive set of features and analyze their contribution to STD individually and as a group. The main goal is to choose a small set of features that are sufficiently informative while keeping the model simple and generalizable. We employ two established models to conduct the analysis: one is linear regression which targets for the most relevant features and the other is logistic linear regression which targets for the most discriminative features. We find the most informative features are comprised of those derived from diverse sources (ASR decoding, duration and lexical properties) and the two models deliver highly consistent feature ranks. STD experiments on both English and Spanish data demonstrate significant performance gains with the proposed feature sets.This work has been partially supported by project PriorSPEECH (TEC2009-14719-C02-01) from the Spanish Ministry of Science and Innovation and by project MAV2VICMR (S2009/TIC-1542) from the Community of Madrid.Elsevier B.V.Departamento de Tecnología Electrónica y de las ComunicacionesEscuela Politécnica SuperiorAnálisis y Tratamiento de Voz y Señales Biométricas (ING EPS-002)Laboratorio de Tecnología Hombre-Computador (ING EPS-010)20142014-09-01research articlehttp://purl.org/coar/resource_type/c_2df8fbb1AMhttp://purl.org/coar/version/c_ab4af688f83e57aainfo:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10486/662781https://dx.doi.org/10.1016/j.csl.2013.09.008reponame:Biblos-e Archivo. Repositorio Institucional de la UAMinstname:Universidad Autónoma de MadridInglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:repositorio.uam.es:10486/6627812026-06-23T12:46:27Z
dc.title.none.fl_str_mv Feature analysis for discriminative confidence estimation in spoken term detection
title Feature analysis for discriminative confidence estimation in spoken term detection
spellingShingle Feature analysis for discriminative confidence estimation in spoken term detection
Tejedor Noguerales, Javier
Discriminative confidence
Feature analysis
Speech recognition
Spoken term detection
Telecomunicaciones
title_short Feature analysis for discriminative confidence estimation in spoken term detection
title_full Feature analysis for discriminative confidence estimation in spoken term detection
title_fullStr Feature analysis for discriminative confidence estimation in spoken term detection
title_full_unstemmed Feature analysis for discriminative confidence estimation in spoken term detection
title_sort Feature analysis for discriminative confidence estimation in spoken term detection
dc.creator.none.fl_str_mv Tejedor Noguerales, Javier
Toledano, Doroteo T.
Wang, Dong
King, Simon
Colás Pasamontes, José
author Tejedor Noguerales, Javier
author_facet Tejedor Noguerales, Javier
Toledano, Doroteo T.
Wang, Dong
King, Simon
Colás Pasamontes, José
author_role author
author2 Toledano, Doroteo T.
Wang, Dong
King, Simon
Colás Pasamontes, José
author2_role author
author
author
author
dc.contributor.none.fl_str_mv Departamento de Tecnología Electrónica y de las Comunicaciones
Escuela Politécnica Superior
Análisis y Tratamiento de Voz y Señales Biométricas (ING EPS-002)
Laboratorio de Tecnología Hombre-Computador (ING EPS-010)
dc.subject.none.fl_str_mv Discriminative confidence
Feature analysis
Speech recognition
Spoken term detection
Telecomunicaciones
topic Discriminative confidence
Feature analysis
Speech recognition
Spoken term detection
Telecomunicaciones
description This is the author’s version of a work that was accepted for publication in Computer Speech & Language. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in Computer Speech & Language, 28, 5, (2014) DOI: 10.1016/j.csl.2013.09.008
publishDate 2014
dc.date.none.fl_str_mv 2014
2014-09-01
dc.type.none.fl_str_mv research article
http://purl.org/coar/resource_type/c_2df8fbb1
AM
http://purl.org/coar/version/c_ab4af688f83e57aa
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv http://hdl.handle.net/10486/662781
https://dx.doi.org/10.1016/j.csl.2013.09.008
url http://hdl.handle.net/10486/662781
https://dx.doi.org/10.1016/j.csl.2013.09.008
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Elsevier B.V.
publisher.none.fl_str_mv Elsevier B.V.
dc.source.none.fl_str_mv reponame:Biblos-e Archivo. Repositorio Institucional de la UAM
instname:Universidad Autónoma de Madrid
instname_str Universidad Autónoma de Madrid
reponame_str Biblos-e Archivo. Repositorio Institucional de la UAM
collection Biblos-e Archivo. Repositorio Institucional de la UAM
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869415484680568832
score 15.198674