Combining semantic and geometric features for object class segmentation of indoor scenes

Scene understanding is a necessary prerequisite for robots acting autonomously in complex environments. Low-cost RGB-D cameras such as Microsoft Kinect enabled new methods for analyzing indoor scenes and are now ubiquitously used in indoor robotics. We investigate strategies for efficient pixelwise...

ver descrição completa

Detalhes bibliográficos
Autores: Husain, Farzad, Schulz, Hannes, Dellen, Babette, Torras, Carme, Behnke, Sven
Formato: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2016
País:España
Recursos:Consejo Superior de Investigaciones Científicas (CSIC)
Repositorio:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/132983
Acesso em linha:http://hdl.handle.net/10261/132983
Access Level:acceso abierto
Palavra-chave:Semantic scene understanding
Categorization
Segmentation
Descrição
Resumo:Scene understanding is a necessary prerequisite for robots acting autonomously in complex environments. Low-cost RGB-D cameras such as Microsoft Kinect enabled new methods for analyzing indoor scenes and are now ubiquitously used in indoor robotics. We investigate strategies for efficient pixelwise object class labeling of indoor scenes that combine both pretrained semantic features transferred from a large color image dataset and geometric features, computed relative to the room structures, including a novel distance-from-wall feature, which encodes the proximity of scene points to a detected major wall of the room. We evaluate our approach on the popular NYU v2 dataset. Several deep learning models are tested, which are designed to exploit different characteristics of the data. This includes feature learning with two different pooling sizes. Our results indicate that combining semantic and geometric features yields significantly improved results for the task of object class segmentation.