Using deep learning for sound classification in citizen science: a practical approach with soundless

The field of Deep Learning has experienced tremendous growth in recent years, sparking interest among users and researchers. However, deploying Deep Learning models in real-world projects presents significant technical challenges. This Master's Thesis provides a practical approach to designing...

Full description

Bibliographic Details
Author: Castelló Tejera, David
Format: master thesis
Publication Date:2023
Country:España
Institution:Universitat Oberta de Catalunya (UOC)
Repository:O2, repositorio institucional de la UOC
OAI Identifier:oai:openaccess.uoc.edu:10609/148916
Online Access:http://hdl.handle.net/10609/148916
Access Level:Open access
Keyword:deep learning
federated learning
TensorFlow
ESC-50
deployment on Android
citizen science
audio classification
Engineering mathematics -- TFM
Matemàtica per a enginyers -- TFM
id ES_9debc7cc6bd17c45d22780b68ffb9f08
oai_identifier_str oai:openaccess.uoc.edu:10609/148916
network_acronym_str ES
network_name_str España
repository_id_str
spelling Using deep learning for sound classification in citizen science: a practical approach with soundlessCastelló Tejera, Daviddeep learningfederated learningTensorFlowESC-50deployment on Androidcitizen scienceaudio classificationEngineering mathematics -- TFMMatemàtica per a enginyers -- TFMThe field of Deep Learning has experienced tremendous growth in recent years, sparking interest among users and researchers. However, deploying Deep Learning models in real-world projects presents significant technical challenges. This Master's Thesis provides a practical approach to designing and constructing a custom Deep Learning model for audio classification, intended for use within the Soundless project—a citizen science platform investigating noise pollution and its impact on human health. The primary objective is to construct a custom model deployable within the Android application of the Soundless project. Different model architectures are explored considering the complexity constraints of deploying Deep Learning models on the edge. The model is built using the TensorFlow framework. Evaluated against the ESC-50 benchmark, the model demonstrates prediction accuracies of over 86%. The model is then integrated into an Android app prototype for testing. A custom dataset is constructed, termed NBAC, comprising 780 audio samples covering 13 distinct classes. NBAC is designed to be aligned with the acoustic context of the Soundless project. The model's performance on NBAC achieves over 90% accuracy. Further, this work investigates various implementation alternatives for utilizing and enhancing the model in a production environment. A centralized improvement approach is proposed, which entails locally storing labeled feature representations of audio samples and training a classifier. Alternatively, a decentralized improvement approach is formulated using Federated Learning. Both strategies, leveraging the custom-designed models, yield promising outcomes. They not only preserve the anticipated accuracies but also facilitate the desired enhancements.Universitat Oberta de Catalunya (UOC)Garcia Lopez, Pedro202320232023info:eu-repo/semantics/masterThesisapplication/pdfapplication/pdfhttp://hdl.handle.net/10609/148916reponame:O2, repositorio institucional de la UOCinstname:Universitat Oberta de Catalunya (UOC)InglésCC BY-NC-NDhttp://creativecommons.org/licenses/by-nc-nd/3.0/es/info:eu-repo/semantics/openAccessoai:openaccess.uoc.edu:10609/1489162026-05-28T12:42:01Z
dc.title.none.fl_str_mv Using deep learning for sound classification in citizen science: a practical approach with soundless
title Using deep learning for sound classification in citizen science: a practical approach with soundless
spellingShingle Using deep learning for sound classification in citizen science: a practical approach with soundless
Castelló Tejera, David
deep learning
federated learning
TensorFlow
ESC-50
deployment on Android
citizen science
audio classification
Engineering mathematics -- TFM
Matemàtica per a enginyers -- TFM
title_short Using deep learning for sound classification in citizen science: a practical approach with soundless
title_full Using deep learning for sound classification in citizen science: a practical approach with soundless
title_fullStr Using deep learning for sound classification in citizen science: a practical approach with soundless
title_full_unstemmed Using deep learning for sound classification in citizen science: a practical approach with soundless
title_sort Using deep learning for sound classification in citizen science: a practical approach with soundless
dc.creator.none.fl_str_mv Castelló Tejera, David
author Castelló Tejera, David
author_facet Castelló Tejera, David
author_role author
dc.contributor.none.fl_str_mv Garcia Lopez, Pedro
dc.subject.none.fl_str_mv deep learning
federated learning
TensorFlow
ESC-50
deployment on Android
citizen science
audio classification
Engineering mathematics -- TFM
Matemàtica per a enginyers -- TFM
topic deep learning
federated learning
TensorFlow
ESC-50
deployment on Android
citizen science
audio classification
Engineering mathematics -- TFM
Matemàtica per a enginyers -- TFM
description The field of Deep Learning has experienced tremendous growth in recent years, sparking interest among users and researchers. However, deploying Deep Learning models in real-world projects presents significant technical challenges. This Master's Thesis provides a practical approach to designing and constructing a custom Deep Learning model for audio classification, intended for use within the Soundless project—a citizen science platform investigating noise pollution and its impact on human health. The primary objective is to construct a custom model deployable within the Android application of the Soundless project. Different model architectures are explored considering the complexity constraints of deploying Deep Learning models on the edge. The model is built using the TensorFlow framework. Evaluated against the ESC-50 benchmark, the model demonstrates prediction accuracies of over 86%. The model is then integrated into an Android app prototype for testing. A custom dataset is constructed, termed NBAC, comprising 780 audio samples covering 13 distinct classes. NBAC is designed to be aligned with the acoustic context of the Soundless project. The model's performance on NBAC achieves over 90% accuracy. Further, this work investigates various implementation alternatives for utilizing and enhancing the model in a production environment. A centralized improvement approach is proposed, which entails locally storing labeled feature representations of audio samples and training a classifier. Alternatively, a decentralized improvement approach is formulated using Federated Learning. Both strategies, leveraging the custom-designed models, yield promising outcomes. They not only preserve the anticipated accuracies but also facilitate the desired enhancements.
publishDate 2023
dc.date.none.fl_str_mv 2023
2023
2023
dc.type.none.fl_str_mv info:eu-repo/semantics/masterThesis
format masterThesis
dc.identifier.none.fl_str_mv http://hdl.handle.net/10609/148916
url http://hdl.handle.net/10609/148916
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.rights.none.fl_str_mv CC BY-NC-ND
http://creativecommons.org/licenses/by-nc-nd/3.0/es/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv CC BY-NC-ND
http://creativecommons.org/licenses/by-nc-nd/3.0/es/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Universitat Oberta de Catalunya (UOC)
publisher.none.fl_str_mv Universitat Oberta de Catalunya (UOC)
dc.source.none.fl_str_mv reponame:O2, repositorio institucional de la UOC
instname:Universitat Oberta de Catalunya (UOC)
instname_str Universitat Oberta de Catalunya (UOC)
reponame_str O2, repositorio institucional de la UOC
collection O2, repositorio institucional de la UOC
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869414782886477824
score 15,301629