Controlling semantics of diffusion-augmented data for unsupervised domain adaptation

Unsupervised domain adaptation (UDA) offers a compelling solution to bridge the gap between labelled synthetic data and unlabelled real‐world data for training semantic segmentation models, given the high costs associated with manual annotation. However, the visual differences between the synthetic...

Full description

Bibliographic Details
Authors: Ridley, Henrietta, Alcover Couso, Roberto, San Miguel Avedillo, Juan Carlos
Format: article
Publication Date:2025
Country:España
Institution:Universidad Autónoma de Madrid
Repository:Biblos-e Archivo. Repositorio Institucional de la UAM
Language:English
OAI Identifier:oai:dnet:biblosearchi::a4c93f36a6fdbb663c7d6a2fee87fc73
Online Access:https://hdl.handle.net/10486/767960
https://dx.doi.org/10.1049/cvi2.70002
Access Level:Open access
Keyword:Computer Vision
Image Segmentation
Unsupervised Learning
Informática
Telecomunicaciones
id ES_0da2af6f1392dbd1b35360ef528a99a5
oai_identifier_str oai:dnet:biblosearchi::a4c93f36a6fdbb663c7d6a2fee87fc73
network_acronym_str ES
network_name_str España
repository_id_str
spelling Controlling semantics of diffusion-augmented data for unsupervised domain adaptationRidley, HenriettaAlcover Couso, RobertoSan Miguel Avedillo, Juan CarlosComputer VisionImage SegmentationUnsupervised LearningInformáticaTelecomunicacionesUnsupervised domain adaptation (UDA) offers a compelling solution to bridge the gap between labelled synthetic data and unlabelled real‐world data for training semantic segmentation models, given the high costs associated with manual annotation. However, the visual differences between the synthetic and real images pose significant challenges to their practical applications. This work addresses these challenges through synthetic‐toreal style transfer leveraging diffusion models. The authors’ proposal incorporates semantic controllers to guide the diffusion process and low‐rank adaptations (LoRAs) to ensure that style‐transferred images align with real‐world aesthetics while preserving semantic layout. Moreover, the authors introduce quality metrics to rank the utility of generated images, enabling the selective use of high‐quality images for training. To further enhance reliability, the authors propose a novel loss function that mitigates artefacts from the style transfer process by incorporating only pixels aligned with the original semantic labels. Experimental results demonstrate that the authors’ proposal outperforms selected state‐of‐the‐art methods for image generation and UDA training, achieving optimal performance even with a smaller set of high‐quality generated images. The authors’ code and models are available at http://www‐vpu.eps.uam.es/ControllingSem4UDA/Agencia Estatal de Investigación de España, Grant/ Award Numbers: PID2021‐125051OB‐I00, TED2021‐131643A‐I00WileyEscuela Politécnica SuperiorDepartamento de Ingeniería InformáticaDepartamento de Tecnología Electrónica y de las ComunicacionesGobierno de España20252025-01-17research articlehttp://purl.org/coar/resource_type/c_2df8fbb1VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfhttps://hdl.handle.net/10486/767960https://dx.doi.org/10.1049/cvi2.70002reponame:Biblos-e Archivo. Repositorio Institucional de la UAMinstname:Universidad Autónoma de MadridInglésengopen accesshttp://purl.org/coar/access_right/c_abf2Attribution 4.0 Internationalhttp://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/openAccessoai:dnet:biblosearchi::a4c93f36a6fdbb663c7d6a2fee87fc732026-06-23T12:46:27Z
dc.title.none.fl_str_mv Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
title Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
spellingShingle Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
Ridley, Henrietta
Computer Vision
Image Segmentation
Unsupervised Learning
Informática
Telecomunicaciones
title_short Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
title_full Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
title_fullStr Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
title_full_unstemmed Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
title_sort Controlling semantics of diffusion-augmented data for unsupervised domain adaptation
dc.creator.none.fl_str_mv Ridley, Henrietta
Alcover Couso, Roberto
San Miguel Avedillo, Juan Carlos
author Ridley, Henrietta
author_facet Ridley, Henrietta
Alcover Couso, Roberto
San Miguel Avedillo, Juan Carlos
author_role author
author2 Alcover Couso, Roberto
San Miguel Avedillo, Juan Carlos
author2_role author
author
dc.contributor.none.fl_str_mv Escuela Politécnica Superior
Departamento de Ingeniería Informática
Departamento de Tecnología Electrónica y de las Comunicaciones
Gobierno de España
dc.subject.none.fl_str_mv Computer Vision
Image Segmentation
Unsupervised Learning
Informática
Telecomunicaciones
topic Computer Vision
Image Segmentation
Unsupervised Learning
Informática
Telecomunicaciones
description Unsupervised domain adaptation (UDA) offers a compelling solution to bridge the gap between labelled synthetic data and unlabelled real‐world data for training semantic segmentation models, given the high costs associated with manual annotation. However, the visual differences between the synthetic and real images pose significant challenges to their practical applications. This work addresses these challenges through synthetic‐toreal style transfer leveraging diffusion models. The authors’ proposal incorporates semantic controllers to guide the diffusion process and low‐rank adaptations (LoRAs) to ensure that style‐transferred images align with real‐world aesthetics while preserving semantic layout. Moreover, the authors introduce quality metrics to rank the utility of generated images, enabling the selective use of high‐quality images for training. To further enhance reliability, the authors propose a novel loss function that mitigates artefacts from the style transfer process by incorporating only pixels aligned with the original semantic labels. Experimental results demonstrate that the authors’ proposal outperforms selected state‐of‐the‐art methods for image generation and UDA training, achieving optimal performance even with a smaller set of high‐quality generated images. The authors’ code and models are available at http://www‐vpu.eps.uam.es/ControllingSem4UDA/
publishDate 2025
dc.date.none.fl_str_mv 2025
2025-01-17
dc.type.none.fl_str_mv research article
http://purl.org/coar/resource_type/c_2df8fbb1
VoR
http://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://hdl.handle.net/10486/767960
https://dx.doi.org/10.1049/cvi2.70002
url https://hdl.handle.net/10486/767960
https://dx.doi.org/10.1049/cvi2.70002
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
Attribution 4.0 International
http://creativecommons.org/licenses/by/4.0/
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
Attribution 4.0 International
http://creativecommons.org/licenses/by/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Wiley
publisher.none.fl_str_mv Wiley
dc.source.none.fl_str_mv reponame:Biblos-e Archivo. Repositorio Institucional de la UAM
instname:Universidad Autónoma de Madrid
instname_str Universidad Autónoma de Madrid
reponame_str Biblos-e Archivo. Repositorio Institucional de la UAM
collection Biblos-e Archivo. Repositorio Institucional de la UAM
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869403351007887360
score 15,812455