RVOS: end-to-end recurrent network for video object segmentation

Multiple object video object segmentation is a challenging task, specially for the zero-shot case, when no object mask is given at the initial frame and the model has to find the objects to be segmented along the sequence. In our work, we propose a Recurrent network for multiple object Video Object...

ver descrição completa

Detalhes bibliográficos
Autores: Ventura Royo, Carles|||0000-0002-3055-339X, Bellver, Míriam, Girbau Xalabarder, Andreu, Salvador Aguilera, Amaia|||0000-0002-9908-1685, Marqués Acosta, Fernando|||0000-0001-8311-1168, Giró Nieto, Xavier|||0000-0002-9935-5332
Tipo de documento: artigo
Data de publicação:2019
País:España
Recursos:Universitat Politècnica de Catalunya (UPC)
Repositório:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglês
OAI Identifier:oai:upcommons.upc.edu:2117/328679
Acesso em linha:https://hdl.handle.net/2117/328679
Access Level:Acceso aberto
Palavra-chave:Image processing -- Digital techniques
Imatges -- Processament -- Tècniques digitals
Àrees temàtiques de la UPC::Enginyeria de la telecomunicació::Processament del senyal::Processament de la imatge i del senyal vídeo
Àrees temàtiques de la UPC::So, imatge i multimèdia::Creació multimèdia::Imatge digital
id ES_fe83bdf52b0f38ec88f5fc5a70a4c83f
oai_identifier_str oai:upcommons.upc.edu:2117/328679
network_acronym_str ES
network_name_str España
repository_id_str
spelling RVOS: end-to-end recurrent network for video object segmentationVentura Royo, Carles|||0000-0002-3055-339XBellver, MíriamGirbau Xalabarder, AndreuSalvador Aguilera, Amaia|||0000-0002-9908-1685Marqués Acosta, Fernando|||0000-0001-8311-1168Giró Nieto, Xavier|||0000-0002-9935-5332Image processing -- Digital techniquesImatges -- Processament -- Tècniques digitalsÀrees temàtiques de la UPC::Enginyeria de la telecomunicació::Processament del senyal::Processament de la imatge i del senyal vídeoÀrees temàtiques de la UPC::So, imatge i multimèdia::Creació multimèdia::Imatge digitalMultiple object video object segmentation is a challenging task, specially for the zero-shot case, when no object mask is given at the initial frame and the model has to find the objects to be segmented along the sequence. In our work, we propose a Recurrent network for multiple object Video Object Segmentation (RVOS) that is fully end-to-end trainable. Our model incorporates recurrence on two different domains: (i) the spatial, which allows to discover the different object instances within a frame, and (ii) the temporal, which allows to keep the coherence of the segmented objects along time. We train RVOS for zero-shot video object segmentation and are the first ones to report quantitative results for DAVIS-2017 and YouTube-VOS benchmarks. Further, we adapt RVOS for one-shot video object segmentation by using the masks obtained in previous time steps as inputs to be processed by the recurrent module. Our model reaches comparable results to state-of-the-art techniques in YouTube-VOS benchmark and outperforms all previous video object segmentation methods not using online learning in the DAVIS-2017 benchmark. Moreover, our model achieves faster inference runtimes than previous methods, reaching 44ms/frame on a P100 GPU.This research was supported by the Spanish Ministry ofEconomy and Competitiveness and the European RegionalDevelopment Fund (TIN2015-66951-C2-2-R, TIN2015-65316-P & TEC2016-75976-R), the BSC-CNS SeveroOchoa SEV-2015-0493 and LaCaixa-Severo Ochoa Inter-national Doctoral Fellowship programs, the 2017 SGR 1414and the Industrial Doctorates 2017-DI-064 & 2017-DI-028from the Government of CataloniaPeer Reviewed20192019-06-1520202020-09-10journal articlehttp://purl.org/coar/resource_type/c_6501VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfhttps://hdl.handle.net/2117/328679reponame:UPCommons. Portal del coneixement obert de la UPCinstname:Universitat Politècnica de Catalunya (UPC)InglésengMinisterio de Economía y Competitividad http://doi.org/10.13039/501100003329 TIN2015-66951-C2-2-R RECONOCIMIENTO VISUAL CON METODOLOGIAS DE APRENDIZAJE DE PRINCIPIO A FIN: MIRANDO LAS PERSONAS Y ENTENDIENDO LAS ESCENASMinisterio de Economía y Competitividad http://doi.org/10.13039/501100003329 TIN2015-65316-P COMPUTACION DE ALTAS PRESTACIONES VIIopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:upcommons.upc.edu:2117/3286792026-05-27T15:37:01Z
dc.title.none.fl_str_mv RVOS: end-to-end recurrent network for video object segmentation
title RVOS: end-to-end recurrent network for video object segmentation
spellingShingle RVOS: end-to-end recurrent network for video object segmentation
Ventura Royo, Carles|||0000-0002-3055-339X
Image processing -- Digital techniques
Imatges -- Processament -- Tècniques digitals
Àrees temàtiques de la UPC::Enginyeria de la telecomunicació::Processament del senyal::Processament de la imatge i del senyal vídeo
Àrees temàtiques de la UPC::So, imatge i multimèdia::Creació multimèdia::Imatge digital
title_short RVOS: end-to-end recurrent network for video object segmentation
title_full RVOS: end-to-end recurrent network for video object segmentation
title_fullStr RVOS: end-to-end recurrent network for video object segmentation
title_full_unstemmed RVOS: end-to-end recurrent network for video object segmentation
title_sort RVOS: end-to-end recurrent network for video object segmentation
dc.creator.none.fl_str_mv Ventura Royo, Carles|||0000-0002-3055-339X
Bellver, Míriam
Girbau Xalabarder, Andreu
Salvador Aguilera, Amaia|||0000-0002-9908-1685
Marqués Acosta, Fernando|||0000-0001-8311-1168
Giró Nieto, Xavier|||0000-0002-9935-5332
author Ventura Royo, Carles|||0000-0002-3055-339X
author_facet Ventura Royo, Carles|||0000-0002-3055-339X
Bellver, Míriam
Girbau Xalabarder, Andreu
Salvador Aguilera, Amaia|||0000-0002-9908-1685
Marqués Acosta, Fernando|||0000-0001-8311-1168
Giró Nieto, Xavier|||0000-0002-9935-5332
author_role author
author2 Bellver, Míriam
Girbau Xalabarder, Andreu
Salvador Aguilera, Amaia|||0000-0002-9908-1685
Marqués Acosta, Fernando|||0000-0001-8311-1168
Giró Nieto, Xavier|||0000-0002-9935-5332
author2_role author
author
author
author
author
dc.subject.none.fl_str_mv Image processing -- Digital techniques
Imatges -- Processament -- Tècniques digitals
Àrees temàtiques de la UPC::Enginyeria de la telecomunicació::Processament del senyal::Processament de la imatge i del senyal vídeo
Àrees temàtiques de la UPC::So, imatge i multimèdia::Creació multimèdia::Imatge digital
topic Image processing -- Digital techniques
Imatges -- Processament -- Tècniques digitals
Àrees temàtiques de la UPC::Enginyeria de la telecomunicació::Processament del senyal::Processament de la imatge i del senyal vídeo
Àrees temàtiques de la UPC::So, imatge i multimèdia::Creació multimèdia::Imatge digital
description Multiple object video object segmentation is a challenging task, specially for the zero-shot case, when no object mask is given at the initial frame and the model has to find the objects to be segmented along the sequence. In our work, we propose a Recurrent network for multiple object Video Object Segmentation (RVOS) that is fully end-to-end trainable. Our model incorporates recurrence on two different domains: (i) the spatial, which allows to discover the different object instances within a frame, and (ii) the temporal, which allows to keep the coherence of the segmented objects along time. We train RVOS for zero-shot video object segmentation and are the first ones to report quantitative results for DAVIS-2017 and YouTube-VOS benchmarks. Further, we adapt RVOS for one-shot video object segmentation by using the masks obtained in previous time steps as inputs to be processed by the recurrent module. Our model reaches comparable results to state-of-the-art techniques in YouTube-VOS benchmark and outperforms all previous video object segmentation methods not using online learning in the DAVIS-2017 benchmark. Moreover, our model achieves faster inference runtimes than previous methods, reaching 44ms/frame on a P100 GPU.
publishDate 2019
dc.date.none.fl_str_mv 2019
2019-06-15
2020
2020-09-10
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
VoR
http://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://hdl.handle.net/2117/328679
url https://hdl.handle.net/2117/328679
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.relation.none.fl_str_mv Ministerio de Economía y Competitividad http://doi.org/10.13039/501100003329 TIN2015-66951-C2-2-R RECONOCIMIENTO VISUAL CON METODOLOGIAS DE APRENDIZAJE DE PRINCIPIO A FIN: MIRANDO LAS PERSONAS Y ENTENDIENDO LAS ESCENAS
Ministerio de Economía y Competitividad http://doi.org/10.13039/501100003329 TIN2015-65316-P COMPUTACION DE ALTAS PRESTACIONES VII
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.source.none.fl_str_mv reponame:UPCommons. Portal del coneixement obert de la UPC
instname:Universitat Politècnica de Catalunya (UPC)
instname_str Universitat Politècnica de Catalunya (UPC)
reponame_str UPCommons. Portal del coneixement obert de la UPC
collection UPCommons. Portal del coneixement obert de la UPC
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869425691964997632
score 15,301629