Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping

[EN] Automatic accent classification is an active research field concerning speech processing. It can be useful to identify a speaker's region of origin, which can be applied in police investigations carried out by Law Enforcement Agencies, as well as for the improvement of current speech recog...

Descripción completa

Detalles Bibliográficos
Autores: Carofilis Vasco, Roberto Andrés, Alegre Gutiérrez, Enrique, Fidalgo Fernández, Eduardo, Fernández Robles, Laura
Tipo de recurso: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2023
País:España
Institución:Universidad de León
Repositorio:BULERIA. Repositorio Institucional de la Universidad de León
OAI Identifier:oai:buleria.unileon.es:10612/23238
Acceso en línea:https://ieeexplore.ieee.org/document/10190103
https://hdl.handle.net/10612/23238
Access Level:acceso abierto
Palabra clave:Informática
Ingeniería de sistemas
Supervised learning
Learning-to-rank
Influence detection
Feature extraction
Darknet
Tor hidden services
3304.05 Sistemas de Reconocimiento de Caracteres
5701.04 Lingüística Informatizada
1203.04 Inteligencia Artificial
1209.03 Análisis de Datos
id ES_cebfebf193f5106a9f975baaedddd0c0
oai_identifier_str oai:buleria.unileon.es:10612/23238
network_acronym_str ES
network_name_str España
repository_id_str
spelling Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation MappingCarofilis Vasco, Roberto AndrésAlegre Gutiérrez, EnriqueFidalgo Fernández, EduardoFernández Robles, LauraInformáticaIngeniería de sistemasSupervised learningLearning-to-rankInfluence detectionFeature extractionDarknetTor hidden services3304.05 Sistemas de Reconocimiento de Caracteres5701.04 Lingüística Informatizada1203.04 Inteligencia Artificial1209.03 Análisis de Datos[EN] Automatic accent classification is an active research field concerning speech processing. It can be useful to identify a speaker's region of origin, which can be applied in police investigations carried out by Law Enforcement Agencies, as well as for the improvement of current speech recognition systems. This article presents a novel descriptor called Grad-Transfer, extracted using the Gradient-weighted Class Activation Mapping (Grad-CAM) method based on convolutional neural network (CNN) interpretability. Additionally, we propose a methodology for accent classification that implements Grad-Transfer, which is based on transferring the knowledge acquired by a CNN to a classical machine learning algorithm. The article works on two hypotheses: the coarse localization maps produced by Grad-CAM on spectrograms are able to highlight the regions of the spectrograms that are important for predicting accents, and Grad-Transfer descriptors computed from audios represent distinctive descriptions of the target accents. These hypotheses were demonstrated experimentally, clustering the generated Grad-Transfer descriptors according to the original accent of the audios using Birch and k -means algorithms. We carried out experiments on the Voice Cloning Toolkit dataset, seeing an increase of macro average accuracy, and unweighted average recall in the results obtained by a Gaussian Naive Bayes classifier up to 23.00%, and 23.58%, respectively, compared to a model trained with spectrograms. This demonstrates that Grad-Transfer is able to improve the performance of accent classification models and opens the door to new implementations in similar tasks.SIThis publication reflects the views only of the authors, and the European Union.s Horizon 2020 Research and Innovation Framework Programme, H2020 SU-FCT-2019 cannot be held responsible for any use which may be made of the information contained therein.European CommissionInstitute of Electrical and Electronics EngineersIngenieria de Sistemas y AutomaticaEscuela de Ingenierias Industrial, Informática y Aeroespacial2023info:eu-repo/semantics/articleinfo:eu-repo/semantics/acceptedVersionhttps://ieeexplore.ieee.org/document/10190103https://hdl.handle.net/10612/23238reponame:BULERIA. Repositorio Institucional de la Universidad de Leóninstname:Universidad de LeónInglésinfo:eu-repo/semantics/openAccessoai:buleria.unileon.es:10612/232382026-06-24T12:43:27Z
dc.title.none.fl_str_mv Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
title Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
spellingShingle Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
Carofilis Vasco, Roberto Andrés
Informática
Ingeniería de sistemas
Supervised learning
Learning-to-rank
Influence detection
Feature extraction
Darknet
Tor hidden services
3304.05 Sistemas de Reconocimiento de Caracteres
5701.04 Lingüística Informatizada
1203.04 Inteligencia Artificial
1209.03 Análisis de Datos
title_short Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
title_full Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
title_fullStr Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
title_full_unstemmed Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
title_sort Improvement of Accent Classification Models Through Grad-Transfer From Spectrograms and Gradient-Weighted Class Activation Mapping
dc.creator.none.fl_str_mv Carofilis Vasco, Roberto Andrés
Alegre Gutiérrez, Enrique
Fidalgo Fernández, Eduardo
Fernández Robles, Laura
author Carofilis Vasco, Roberto Andrés
author_facet Carofilis Vasco, Roberto Andrés
Alegre Gutiérrez, Enrique
Fidalgo Fernández, Eduardo
Fernández Robles, Laura
author_role author
author2 Alegre Gutiérrez, Enrique
Fidalgo Fernández, Eduardo
Fernández Robles, Laura
author2_role author
author
author
dc.contributor.none.fl_str_mv Ingenieria de Sistemas y Automatica
Escuela de Ingenierias Industrial, Informática y Aeroespacial
dc.subject.none.fl_str_mv Informática
Ingeniería de sistemas
Supervised learning
Learning-to-rank
Influence detection
Feature extraction
Darknet
Tor hidden services
3304.05 Sistemas de Reconocimiento de Caracteres
5701.04 Lingüística Informatizada
1203.04 Inteligencia Artificial
1209.03 Análisis de Datos
topic Informática
Ingeniería de sistemas
Supervised learning
Learning-to-rank
Influence detection
Feature extraction
Darknet
Tor hidden services
3304.05 Sistemas de Reconocimiento de Caracteres
5701.04 Lingüística Informatizada
1203.04 Inteligencia Artificial
1209.03 Análisis de Datos
description [EN] Automatic accent classification is an active research field concerning speech processing. It can be useful to identify a speaker's region of origin, which can be applied in police investigations carried out by Law Enforcement Agencies, as well as for the improvement of current speech recognition systems. This article presents a novel descriptor called Grad-Transfer, extracted using the Gradient-weighted Class Activation Mapping (Grad-CAM) method based on convolutional neural network (CNN) interpretability. Additionally, we propose a methodology for accent classification that implements Grad-Transfer, which is based on transferring the knowledge acquired by a CNN to a classical machine learning algorithm. The article works on two hypotheses: the coarse localization maps produced by Grad-CAM on spectrograms are able to highlight the regions of the spectrograms that are important for predicting accents, and Grad-Transfer descriptors computed from audios represent distinctive descriptions of the target accents. These hypotheses were demonstrated experimentally, clustering the generated Grad-Transfer descriptors according to the original accent of the audios using Birch and k -means algorithms. We carried out experiments on the Voice Cloning Toolkit dataset, seeing an increase of macro average accuracy, and unweighted average recall in the results obtained by a Gaussian Naive Bayes classifier up to 23.00%, and 23.58%, respectively, compared to a model trained with spectrograms. This demonstrates that Grad-Transfer is able to improve the performance of accent classification models and opens the door to new implementations in similar tasks.
publishDate 2023
dc.date.none.fl_str_mv 2023
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/acceptedVersion
format article
status_str acceptedVersion
dc.identifier.none.fl_str_mv https://ieeexplore.ieee.org/document/10190103
https://hdl.handle.net/10612/23238
url https://ieeexplore.ieee.org/document/10190103
https://hdl.handle.net/10612/23238
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
eu_rights_str_mv openAccess
dc.publisher.none.fl_str_mv Institute of Electrical and Electronics Engineers
publisher.none.fl_str_mv Institute of Electrical and Electronics Engineers
dc.source.none.fl_str_mv reponame:BULERIA. Repositorio Institucional de la Universidad de León
instname:Universidad de León
instname_str Universidad de León
reponame_str BULERIA. Repositorio Institucional de la Universidad de León
collection BULERIA. Repositorio Institucional de la Universidad de León
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869420019010502656
score 15.812455