Learning the semantics of object-action relations by observation

Recognizing manipulations performed by a human and the transfer and execution of this by a robot is a difficult problem. We address this in the current study by introducing a novel representation of the relations between objects at decisive time points during a manipulation. Thereby, we encode the e...

Descripción completa

Detalles Bibliográficos
Autores: Aksoy, Eren Erdal, Abramov, Alexey, Dörr, Johannes, Ning, Kejun, Dellen, Babette, Wörgötter, Florentin
Tipo de recurso: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2011
País:España
Institución:Consejo Superior de Investigaciones Científicas (CSIC)
Repositorio:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/96813
Acceso en línea:http://hdl.handle.net/10261/96813
Access Level:acceso abierto
Palabra clave:Action recognition
Affordances
Object categorization
Semantic scene graphs Unsupervised learning
Object–action complexes (OACs)
id ES_a992afd53d4a7c7c8ceb4b661481ab13
oai_identifier_str oai:digital.csic.es:10261/96813
network_acronym_str ES
network_name_str España
repository_id_str
spelling Learning the semantics of object-action relations by observationAksoy, Eren ErdalAbramov, AlexeyDörr, JohannesNing, KejunDellen, BabetteWörgötter, FlorentinAction recognitionAffordancesObject categorizationSemantic scene graphs Unsupervised learningObject–action complexes (OACs)Recognizing manipulations performed by a human and the transfer and execution of this by a robot is a difficult problem. We address this in the current study by introducing a novel representation of the relations between objects at decisive time points during a manipulation. Thereby, we encode the essential changes in a visual scenery in a condensed way such that a robot can recognize and learn a manipulation without prior object knowledge. To achieve this we continuously track image segments in the video and construct a dynamic graph sequence. Topological transitions of those graphs occur whenever a spatial relation between some segments has changed in a discontinuous way and these moments are stored in a transition matrix called the semantic event chain (SEC). We demonstrate that these time points are highly descriptive for distinguishing between different manipulations. Employing simple sub-string search algorithms, SECs can be compared and type-similar manipulations can be recognized with high confidence. As the approach is generic, statistical learning can be used to find the archetypal SEC of a given manipulation class. The performance of the algorithm is demonstrated on a set of real videos showing hands manipulating various objects and performing different actions. In experiments with a robotic arm, we show that the SEC can be learned by observing human manipulations, transferred to a new scenario, and then reproduced by the machine. © SAGE Publications 2011.The research leading to these results has received funding from the European Community's Seventh Framework Programme FP7=2007-2013 - Challenge 2 - Cognitive Systems, Interaction, Robotics - under grant agreement No 247947 - GARNICS. B.D. acknowledges support from the Spanish Ministry for Science and Innovation via a Ramon y Cajal Fellowship.Peer ReviewedSage Publications2014201420112014info:eu-repo/semantics/articlehttp://purl.org/coar/resource_type/c_6501Postprintinfo:eu-repo/semantics/acceptedVersionhttp://hdl.handle.net/10261/96813reponame:DIGITAL.CSIC. Repositorio Institucional del CSICinstname:Consejo Superior de Investigaciones Científicas (CSIC)Inglés#PLACEHOLDER_PARENT_METADATA_VALUE#info:eu-repo/grantAgreement/EC/FP7/247947http://dx.doi.org/10.1177/0278364911410459info:eu-repo/semantics/openAccessoai:digital.csic.es:10261/968132026-05-22T06:33:51Z
dc.title.none.fl_str_mv Learning the semantics of object-action relations by observation
title Learning the semantics of object-action relations by observation
spellingShingle Learning the semantics of object-action relations by observation
Aksoy, Eren Erdal
Action recognition
Affordances
Object categorization
Semantic scene graphs Unsupervised learning
Object–action complexes (OACs)
title_short Learning the semantics of object-action relations by observation
title_full Learning the semantics of object-action relations by observation
title_fullStr Learning the semantics of object-action relations by observation
title_full_unstemmed Learning the semantics of object-action relations by observation
title_sort Learning the semantics of object-action relations by observation
dc.creator.none.fl_str_mv Aksoy, Eren Erdal
Abramov, Alexey
Dörr, Johannes
Ning, Kejun
Dellen, Babette
Wörgötter, Florentin
author Aksoy, Eren Erdal
author_facet Aksoy, Eren Erdal
Abramov, Alexey
Dörr, Johannes
Ning, Kejun
Dellen, Babette
Wörgötter, Florentin
author_role author
author2 Abramov, Alexey
Dörr, Johannes
Ning, Kejun
Dellen, Babette
Wörgötter, Florentin
author2_role author
author
author
author
author
dc.subject.none.fl_str_mv Action recognition
Affordances
Object categorization
Semantic scene graphs Unsupervised learning
Object–action complexes (OACs)
topic Action recognition
Affordances
Object categorization
Semantic scene graphs Unsupervised learning
Object–action complexes (OACs)
description Recognizing manipulations performed by a human and the transfer and execution of this by a robot is a difficult problem. We address this in the current study by introducing a novel representation of the relations between objects at decisive time points during a manipulation. Thereby, we encode the essential changes in a visual scenery in a condensed way such that a robot can recognize and learn a manipulation without prior object knowledge. To achieve this we continuously track image segments in the video and construct a dynamic graph sequence. Topological transitions of those graphs occur whenever a spatial relation between some segments has changed in a discontinuous way and these moments are stored in a transition matrix called the semantic event chain (SEC). We demonstrate that these time points are highly descriptive for distinguishing between different manipulations. Employing simple sub-string search algorithms, SECs can be compared and type-similar manipulations can be recognized with high confidence. As the approach is generic, statistical learning can be used to find the archetypal SEC of a given manipulation class. The performance of the algorithm is demonstrated on a set of real videos showing hands manipulating various objects and performing different actions. In experiments with a robotic arm, we show that the SEC can be learned by observing human manipulations, transferred to a new scenario, and then reproduced by the machine. © SAGE Publications 2011.
publishDate 2011
dc.date.none.fl_str_mv 2011
2014
2014
2014
dc.type.none.fl_str_mv info:eu-repo/semantics/article
http://purl.org/coar/resource_type/c_6501
Postprint
info:eu-repo/semantics/acceptedVersion
format article
status_str acceptedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10261/96813
url http://hdl.handle.net/10261/96813
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv #PLACEHOLDER_PARENT_METADATA_VALUE#
info:eu-repo/grantAgreement/EC/FP7/247947
http://dx.doi.org/10.1177/0278364911410459
dc.rights.none.fl_str_mv info:eu-repo/semantics/openAccess
eu_rights_str_mv openAccess
dc.publisher.none.fl_str_mv Sage Publications
publisher.none.fl_str_mv Sage Publications
dc.source.none.fl_str_mv reponame:DIGITAL.CSIC. Repositorio Institucional del CSIC
instname:Consejo Superior de Investigaciones Científicas (CSIC)
instname_str Consejo Superior de Investigaciones Científicas (CSIC)
reponame_str DIGITAL.CSIC. Repositorio Institucional del CSIC
collection DIGITAL.CSIC. Repositorio Institucional del CSIC
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869416033237860352
score 15,812455