Feature analysis for discriminative confidence estimation in spoken term detection
This is the author’s version of a work that was accepted for publication in Computer Speech & Language. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Cha...
| Autores: | , , , , |
|---|---|
| Tipo de recurso: | artículo |
| Fecha de publicación: | 2014 |
| País: | España |
| Institución: | Universidad Autónoma de Madrid |
| Repositorio: | Biblos-e Archivo. Repositorio Institucional de la UAM |
| Idioma: | inglés |
| OAI Identifier: | oai:repositorio.uam.es:10486/662781 |
| Acceso en línea: | http://hdl.handle.net/10486/662781 https://dx.doi.org/10.1016/j.csl.2013.09.008 |
| Access Level: | acceso abierto |
| Palabra clave: | Discriminative confidence Feature analysis Speech recognition Spoken term detection Telecomunicaciones |
| id |
ES_a448de7a2b255785d5b186a80002fbf6 |
|---|---|
| oai_identifier_str |
oai:repositorio.uam.es:10486/662781 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Feature analysis for discriminative confidence estimation in spoken term detectionTejedor Noguerales, JavierToledano, Doroteo T.Wang, DongKing, SimonColás Pasamontes, JoséDiscriminative confidenceFeature analysisSpeech recognitionSpoken term detectionTelecomunicacionesThis is the author’s version of a work that was accepted for publication in Computer Speech & Language. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in Computer Speech & Language, 28, 5, (2014) DOI: 10.1016/j.csl.2013.09.008Discriminative confidence based on multi-layer perceptrons (MLPs) and multiple features has shown significant advantage compared to the widely used lattice-based confidence in spoken term detection (STD). Although the MLP-based framework can handle any features derived from a multitude of sources, choosing all possible features may lead to over complex models and hence less generality. In this paper, we design an extensive set of features and analyze their contribution to STD individually and as a group. The main goal is to choose a small set of features that are sufficiently informative while keeping the model simple and generalizable. We employ two established models to conduct the analysis: one is linear regression which targets for the most relevant features and the other is logistic linear regression which targets for the most discriminative features. We find the most informative features are comprised of those derived from diverse sources (ASR decoding, duration and lexical properties) and the two models deliver highly consistent feature ranks. STD experiments on both English and Spanish data demonstrate significant performance gains with the proposed feature sets.This work has been partially supported by project PriorSPEECH (TEC2009-14719-C02-01) from the Spanish Ministry of Science and Innovation and by project MAV2VICMR (S2009/TIC-1542) from the Community of Madrid.Elsevier B.V.Departamento de Tecnología Electrónica y de las ComunicacionesEscuela Politécnica SuperiorAnálisis y Tratamiento de Voz y Señales Biométricas (ING EPS-002)Laboratorio de Tecnología Hombre-Computador (ING EPS-010)20142014-09-01research articlehttp://purl.org/coar/resource_type/c_2df8fbb1AMhttp://purl.org/coar/version/c_ab4af688f83e57aainfo:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10486/662781https://dx.doi.org/10.1016/j.csl.2013.09.008reponame:Biblos-e Archivo. Repositorio Institucional de la UAMinstname:Universidad Autónoma de MadridInglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:repositorio.uam.es:10486/6627812026-06-23T12:46:27Z |
| dc.title.none.fl_str_mv |
Feature analysis for discriminative confidence estimation in spoken term detection |
| title |
Feature analysis for discriminative confidence estimation in spoken term detection |
| spellingShingle |
Feature analysis for discriminative confidence estimation in spoken term detection Tejedor Noguerales, Javier Discriminative confidence Feature analysis Speech recognition Spoken term detection Telecomunicaciones |
| title_short |
Feature analysis for discriminative confidence estimation in spoken term detection |
| title_full |
Feature analysis for discriminative confidence estimation in spoken term detection |
| title_fullStr |
Feature analysis for discriminative confidence estimation in spoken term detection |
| title_full_unstemmed |
Feature analysis for discriminative confidence estimation in spoken term detection |
| title_sort |
Feature analysis for discriminative confidence estimation in spoken term detection |
| dc.creator.none.fl_str_mv |
Tejedor Noguerales, Javier Toledano, Doroteo T. Wang, Dong King, Simon Colás Pasamontes, José |
| author |
Tejedor Noguerales, Javier |
| author_facet |
Tejedor Noguerales, Javier Toledano, Doroteo T. Wang, Dong King, Simon Colás Pasamontes, José |
| author_role |
author |
| author2 |
Toledano, Doroteo T. Wang, Dong King, Simon Colás Pasamontes, José |
| author2_role |
author author author author |
| dc.contributor.none.fl_str_mv |
Departamento de Tecnología Electrónica y de las Comunicaciones Escuela Politécnica Superior Análisis y Tratamiento de Voz y Señales Biométricas (ING EPS-002) Laboratorio de Tecnología Hombre-Computador (ING EPS-010) |
| dc.subject.none.fl_str_mv |
Discriminative confidence Feature analysis Speech recognition Spoken term detection Telecomunicaciones |
| topic |
Discriminative confidence Feature analysis Speech recognition Spoken term detection Telecomunicaciones |
| description |
This is the author’s version of a work that was accepted for publication in Computer Speech & Language. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in Computer Speech & Language, 28, 5, (2014) DOI: 10.1016/j.csl.2013.09.008 |
| publishDate |
2014 |
| dc.date.none.fl_str_mv |
2014 2014-09-01 |
| dc.type.none.fl_str_mv |
research article http://purl.org/coar/resource_type/c_2df8fbb1 AM http://purl.org/coar/version/c_ab4af688f83e57aa |
| dc.type.openaire.fl_str_mv |
info:eu-repo/semantics/article |
| format |
article |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10486/662781 https://dx.doi.org/10.1016/j.csl.2013.09.008 |
| url |
http://hdl.handle.net/10486/662781 https://dx.doi.org/10.1016/j.csl.2013.09.008 |
| dc.language.none.fl_str_mv |
Inglés eng |
| language_invalid_str_mv |
Inglés |
| language |
eng |
| dc.rights.none.fl_str_mv |
open access http://purl.org/coar/access_right/c_abf2 |
| dc.rights.openaire.fl_str_mv |
info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
open access http://purl.org/coar/access_right/c_abf2 |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf |
| dc.publisher.none.fl_str_mv |
Elsevier B.V. |
| publisher.none.fl_str_mv |
Elsevier B.V. |
| dc.source.none.fl_str_mv |
reponame:Biblos-e Archivo. Repositorio Institucional de la UAM instname:Universidad Autónoma de Madrid |
| instname_str |
Universidad Autónoma de Madrid |
| reponame_str |
Biblos-e Archivo. Repositorio Institucional de la UAM |
| collection |
Biblos-e Archivo. Repositorio Institucional de la UAM |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869415484680568832 |
| score |
15.198674 |