A comparison of grapheme and phoneme-based units for Spanish spoken term detection

This manuscript version is made available under the CC-BY-NC-ND 4.0 licence http://creativecommons.org/licenses/by-nc-nd/4.0/

Detalhes bibliográficos
Autores: Tejedor Noguerales, Javier, Wang, Dong, Frankel, Joe, King, Simon, Colás Pasamontes, José
Formato: artículo
Fecha de publicación:2008
País:España
Recursos:Universidad Autónoma de Madrid
Repositorio:Biblos-e Archivo. Repositorio Institucional de la UAM
Idioma:inglés
OAI Identifier:oai:repositorio.uam.es:10486/720707
Acesso em linha:http://hdl.handle.net/10486/720707
https://dx.doi.org/10.1016/j.specom.2008.03.005
Access Level:acceso abierto
Palavra-chave:Graphemes
Keyword spotting
Spanish
Spoken term detection
Telecomunicaciones
id ES_17a93d02c681b5d62fe91c6f86d8404d
oai_identifier_str oai:repositorio.uam.es:10486/720707
network_acronym_str ES
network_name_str España
repository_id_str
spelling A comparison of grapheme and phoneme-based units for Spanish spoken term detectionTejedor Noguerales, JavierWang, DongFrankel, JoeKing, SimonColás Pasamontes, JoséGraphemesKeyword spottingSpanishSpoken term detectionTelecomunicacionesThis manuscript version is made available under the CC-BY-NC-ND 4.0 licence http://creativecommons.org/licenses/by-nc-nd/4.0/The ever-increasing volume of audio data available online through the world wide web means that automatic methods for indexing and search are becoming essential. Hidden Markov model (HMM) keyword spotting and lattice search techniques are the two most common approaches used by such systems. In keyword spotting, models or templates are defined for each search term prior to accessing the speech and used to find matches. Lattice search (referred to as spoken term detection), uses a pre-indexing of speech data in terms of word or sub-word units, which can then quickly be searched for arbitrary terms without referring to the original audio. In both cases, the search term can be modelled in terms of sub-word units, typically phonemes. For in-vocabulary words (i.e. words that appear in the pronunciation dictionary), the letter-to-sound conversion systems are accepted to work well. However, for out-of-vocabulary (OOV) search terms, letter-to-sound conversion must be used to generate a pronunciation for the search term. This is usually a hard decision (i.e. not probabilistic and with no possibility of backtracking), and errors introduced at this step are difficult to recover from. We therefore propose the direct use of graphemes (i.e., letter-based sub-word units) for acoustic modelling. This is expected to work particularly well in languages such as Spanish, where despite the letter-to-sound mapping being very regular, the correspondence is not one-to-one, and there will be benefits from avoiding hard decisions at early stages of processing. In this article, we compare three approaches for Spanish keyword spotting or spoken term detection, and within each of these we compare acoustic modelling based on phone and grapheme units. Experiments were performed using the Spanish geographical-domain Albayzin corpus. Results achieved in the two approaches proposed for spoken term detection show us that trigrapheme units for acoustic modelling match or exceed the performance of phone-based acoustic models. In the method proposed for keyword spotting, the results achieved with each acoustic model are very similarThis work was partly funded by the Spanish Ministry of Science and Education (TIN 2005-06885). DW is a Fellow on the Edinburgh Speech Science and Technology (EdSST) interdisciplinary Marie Curie training programme. JF is funded by Scottish Enterprise under the Edinburgh Stanford Link. SK is an EPSRC Advanced Research FellowElsevierDepartamento de Tecnología Electrónica y de las ComunicacionesEscuela Politécnica Superior20082008-03-28research articlehttp://purl.org/coar/resource_type/c_2df8fbb1AMhttp://purl.org/coar/version/c_ab4af688f83e57aainfo:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10486/720707https://dx.doi.org/10.1016/j.specom.2008.03.005reponame:Biblos-e Archivo. Repositorio Institucional de la UAMinstname:Universidad Autónoma de MadridInglésengopen accesshttp://purl.org/coar/access_right/c_abf2Attribution-NonCommercial-NoDerivatives 4.0 Internationalhttp://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccessoai:repositorio.uam.es:10486/7207072026-06-23T12:46:27Z
dc.title.none.fl_str_mv A comparison of grapheme and phoneme-based units for Spanish spoken term detection
title A comparison of grapheme and phoneme-based units for Spanish spoken term detection
spellingShingle A comparison of grapheme and phoneme-based units for Spanish spoken term detection
Tejedor Noguerales, Javier
Graphemes
Keyword spotting
Spanish
Spoken term detection
Telecomunicaciones
title_short A comparison of grapheme and phoneme-based units for Spanish spoken term detection
title_full A comparison of grapheme and phoneme-based units for Spanish spoken term detection
title_fullStr A comparison of grapheme and phoneme-based units for Spanish spoken term detection
title_full_unstemmed A comparison of grapheme and phoneme-based units for Spanish spoken term detection
title_sort A comparison of grapheme and phoneme-based units for Spanish spoken term detection
dc.creator.none.fl_str_mv Tejedor Noguerales, Javier
Wang, Dong
Frankel, Joe
King, Simon
Colás Pasamontes, José
author Tejedor Noguerales, Javier
author_facet Tejedor Noguerales, Javier
Wang, Dong
Frankel, Joe
King, Simon
Colás Pasamontes, José
author_role author
author2 Wang, Dong
Frankel, Joe
King, Simon
Colás Pasamontes, José
author2_role author
author
author
author
dc.contributor.none.fl_str_mv Departamento de Tecnología Electrónica y de las Comunicaciones
Escuela Politécnica Superior
dc.subject.none.fl_str_mv Graphemes
Keyword spotting
Spanish
Spoken term detection
Telecomunicaciones
topic Graphemes
Keyword spotting
Spanish
Spoken term detection
Telecomunicaciones
description This manuscript version is made available under the CC-BY-NC-ND 4.0 licence http://creativecommons.org/licenses/by-nc-nd/4.0/
publishDate 2008
dc.date.none.fl_str_mv 2008
2008-03-28
dc.type.none.fl_str_mv research article
http://purl.org/coar/resource_type/c_2df8fbb1
AM
http://purl.org/coar/version/c_ab4af688f83e57aa
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv http://hdl.handle.net/10486/720707
https://dx.doi.org/10.1016/j.specom.2008.03.005
url http://hdl.handle.net/10486/720707
https://dx.doi.org/10.1016/j.specom.2008.03.005
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
Attribution-NonCommercial-NoDerivatives 4.0 International
http://creativecommons.org/licenses/by-nc-nd/4.0/
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
Attribution-NonCommercial-NoDerivatives 4.0 International
http://creativecommons.org/licenses/by-nc-nd/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:Biblos-e Archivo. Repositorio Institucional de la UAM
instname:Universidad Autónoma de Madrid
instname_str Universidad Autónoma de Madrid
reponame_str Biblos-e Archivo. Repositorio Institucional de la UAM
collection Biblos-e Archivo. Repositorio Institucional de la UAM
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869403935694913536
score 15,223283