Combining semantic and geometric features for object class segmentation of indoor scenes

Scene understanding is a necessary prerequisite for robots acting autonomously in complex environments. Low-cost RGB-D cameras such as Microsoft Kinect enabled new methods for analyzing indoor scenes and are now ubiquitously used in indoor robotics. We investigate strategies for efficient pixelwise...

Descripción completa

Detalles Bibliográficos
Autores: Husain, Farzad, Schulz, Hannes, Dellen, Babette, Torras, Carme, Behnke, Sven
Tipo de recurso: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2016
País:España
Institución:Consejo Superior de Investigaciones Científicas (CSIC)
Repositorio:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/132983
Acceso en línea:http://hdl.handle.net/10261/132983
Access Level:acceso abierto
Palabra clave:Semantic scene understanding
Categorization
Segmentation
Descripción
Sumario:Scene understanding is a necessary prerequisite for robots acting autonomously in complex environments. Low-cost RGB-D cameras such as Microsoft Kinect enabled new methods for analyzing indoor scenes and are now ubiquitously used in indoor robotics. We investigate strategies for efficient pixelwise object class labeling of indoor scenes that combine both pretrained semantic features transferred from a large color image dataset and geometric features, computed relative to the room structures, including a novel distance-from-wall feature, which encodes the proximity of scene points to a detected major wall of the room. We evaluate our approach on the popular NYU v2 dataset. Several deep learning models are tested, which are designed to exploit different characteristics of the data. This includes feature learning with two different pooling sizes. Our results indicate that combining semantic and geometric features yields significantly improved results for the task of object class segmentation.