Decoding class dynamics in learning with noisy labels

The creation of large-scale datasets annotated by humans inevitably introduces noisy labels, leading to reduced generalization in deep-learning models. Sample selection-based learning with noisy labels is a recent approach that exhibits promising upbeat performance improvements. The selection of cle...

Descripción completa

Detalles Bibliográficos
Autores: Tatjer, Albert, Nagarajan, Bhalaji, Marques, Ricardo, Radeva, Petia
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2024
País:España
Institución:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
Repositorio:Recercat. Dipósit de la Recerca de Catalunya
OAI Identifier:oai:recercat.cat:10230/71851
Acceso en línea:http://hdl.handle.net/10230/71851
http://dx.doi.org/10.1016/j.patrec.2024.04.012
Access Level:acceso abierto
Palabra clave:Learning with noisy labels
Label noise modelling
Class dynamics
id ES_8d272bfe5444de99abaeb2d0bbd56a6d
oai_identifier_str oai:recercat.cat:10230/71851
network_acronym_str ES
network_name_str España
repository_id_str
spelling Decoding class dynamics in learning with noisy labelsTatjer, AlbertNagarajan, BhalajiMarques, RicardoRadeva, PetiaLearning with noisy labelsLabel noise modellingClass dynamicsThe creation of large-scale datasets annotated by humans inevitably introduces noisy labels, leading to reduced generalization in deep-learning models. Sample selection-based learning with noisy labels is a recent approach that exhibits promising upbeat performance improvements. The selection of clean samples amongst the noisy samples is an important criterion in the learning process of these models. In this work, we delve deeper into the clean-noise split decision and highlight the aspect that effective demarcation of samples would lead to better performance. We identify the Global Noise Conundrum in the existing models, where the distribution of samples is treated globally. We propose a per-class-based local distribution of samples and demonstrate the effectiveness of this approach in having a better clean-noise split. We validate our proposal on several benchmarks — both real and synthetic, and show substantial improvements over different state-of-the-art algorithms. We further propose a new metric, classiness to extend our analysis and highlight the effectiveness of the proposed method. Source code and instructions to reproduce this paper are available at https://github.com/aldakata/CCLM/This work was partially funded by the EU project MUSAE (No. 01070421), 2021-SGR-01094 (AGAUR), Icrea Academia’2022 (Generalitat de Catalunya), Robo STEAM (2022-1-BG01-KA220-VET-000089434, Erasmus+ EU), DeepSense (ACE053/22/000029, ACCIÓ), CERCA Programme/Generalitat de Catalunya, and Grants PID2022-141566NB-I00 (IDEATE), PDC2022-133642-I00 (DeepFoodVol), and CNS2022-135480 (A-BMC) funded by MICIU/AEI/10.13039/501100011033, by FEDER (UE), and by European Union NextGenerationEU/ PRTR. R. Marques acknowledges the support of the Serra Húnter Programme. B. Nagarajan acknowledges the support of FPI Becas, MICINN, Spain .Elsevier202520252024info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfapplication/pdfhttp://hdl.handle.net/10230/71851http://dx.doi.org/10.1016/j.patrec.2024.04.012http://hdl.handle.net/10230/71851reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésPattern Recognition Letters. 2024 Aug;184:239-45info:eu-repo/grantAgreement/EC/H2020/01070421info:eu-repo/grantAgreement/ES/3PE/PID2022-141566NB-I00info:eu-repo/grantAgreement/ES/3PE/PDC2022-133642-I00info:eu-repo/grantAgreement/ES/3PE/CNS2022-135480© 2024 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).http://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/openAccessoai:recercat.cat:10230/718512026-05-29T05:05:01Z
dc.title.none.fl_str_mv Decoding class dynamics in learning with noisy labels
title Decoding class dynamics in learning with noisy labels
spellingShingle Decoding class dynamics in learning with noisy labels
Tatjer, Albert
Learning with noisy labels
Label noise modelling
Class dynamics
title_short Decoding class dynamics in learning with noisy labels
title_full Decoding class dynamics in learning with noisy labels
title_fullStr Decoding class dynamics in learning with noisy labels
title_full_unstemmed Decoding class dynamics in learning with noisy labels
title_sort Decoding class dynamics in learning with noisy labels
dc.creator.none.fl_str_mv Tatjer, Albert
Nagarajan, Bhalaji
Marques, Ricardo
Radeva, Petia
author Tatjer, Albert
author_facet Tatjer, Albert
Nagarajan, Bhalaji
Marques, Ricardo
Radeva, Petia
author_role author
author2 Nagarajan, Bhalaji
Marques, Ricardo
Radeva, Petia
author2_role author
author
author
dc.subject.none.fl_str_mv Learning with noisy labels
Label noise modelling
Class dynamics
topic Learning with noisy labels
Label noise modelling
Class dynamics
description The creation of large-scale datasets annotated by humans inevitably introduces noisy labels, leading to reduced generalization in deep-learning models. Sample selection-based learning with noisy labels is a recent approach that exhibits promising upbeat performance improvements. The selection of clean samples amongst the noisy samples is an important criterion in the learning process of these models. In this work, we delve deeper into the clean-noise split decision and highlight the aspect that effective demarcation of samples would lead to better performance. We identify the Global Noise Conundrum in the existing models, where the distribution of samples is treated globally. We propose a per-class-based local distribution of samples and demonstrate the effectiveness of this approach in having a better clean-noise split. We validate our proposal on several benchmarks — both real and synthetic, and show substantial improvements over different state-of-the-art algorithms. We further propose a new metric, classiness to extend our analysis and highlight the effectiveness of the proposed method. Source code and instructions to reproduce this paper are available at https://github.com/aldakata/CCLM/
publishDate 2024
dc.date.none.fl_str_mv 2024
2025
2025
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10230/71851
http://dx.doi.org/10.1016/j.patrec.2024.04.012
http://hdl.handle.net/10230/71851
url http://hdl.handle.net/10230/71851
http://dx.doi.org/10.1016/j.patrec.2024.04.012
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Pattern Recognition Letters. 2024 Aug;184:239-45
info:eu-repo/grantAgreement/EC/H2020/01070421
info:eu-repo/grantAgreement/ES/3PE/PID2022-141566NB-I00
info:eu-repo/grantAgreement/ES/3PE/PDC2022-133642-I00
info:eu-repo/grantAgreement/ES/3PE/CNS2022-135480
dc.rights.none.fl_str_mv http://creativecommons.org/licenses/by/4.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv http://creativecommons.org/licenses/by/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:Recercat. Dipósit de la Recerca de Catalunya
instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
instname_str Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
reponame_str Recercat. Dipósit de la Recerca de Catalunya
collection Recercat. Dipósit de la Recerca de Catalunya
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869413020993585152
score 15.228081