Learning the semantics of object-action relations by observation
Recognizing manipulations performed by a human and the transfer and execution of this by a robot is a difficult problem. We address this in the current study by introducing a novel representation of the relations between objects at decisive time points during a manipulation. Thereby, we encode the e...
| Autores: | , , , , , |
|---|---|
| Tipo de recurso: | artículo |
| Estado: | Versión aceptada para publicación |
| Fecha de publicación: | 2011 |
| País: | España |
| Institución: | Consejo Superior de Investigaciones Científicas (CSIC) |
| Repositorio: | DIGITAL.CSIC. Repositorio Institucional del CSIC |
| OAI Identifier: | oai:digital.csic.es:10261/96813 |
| Acceso en línea: | http://hdl.handle.net/10261/96813 |
| Access Level: | acceso abierto |
| Palabra clave: | Action recognition Affordances Object categorization Semantic scene graphs Unsupervised learning Object–action complexes (OACs) |
| id |
ES_a992afd53d4a7c7c8ceb4b661481ab13 |
|---|---|
| oai_identifier_str |
oai:digital.csic.es:10261/96813 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Learning the semantics of object-action relations by observationAksoy, Eren ErdalAbramov, AlexeyDörr, JohannesNing, KejunDellen, BabetteWörgötter, FlorentinAction recognitionAffordancesObject categorizationSemantic scene graphs Unsupervised learningObject–action complexes (OACs)Recognizing manipulations performed by a human and the transfer and execution of this by a robot is a difficult problem. We address this in the current study by introducing a novel representation of the relations between objects at decisive time points during a manipulation. Thereby, we encode the essential changes in a visual scenery in a condensed way such that a robot can recognize and learn a manipulation without prior object knowledge. To achieve this we continuously track image segments in the video and construct a dynamic graph sequence. Topological transitions of those graphs occur whenever a spatial relation between some segments has changed in a discontinuous way and these moments are stored in a transition matrix called the semantic event chain (SEC). We demonstrate that these time points are highly descriptive for distinguishing between different manipulations. Employing simple sub-string search algorithms, SECs can be compared and type-similar manipulations can be recognized with high confidence. As the approach is generic, statistical learning can be used to find the archetypal SEC of a given manipulation class. The performance of the algorithm is demonstrated on a set of real videos showing hands manipulating various objects and performing different actions. In experiments with a robotic arm, we show that the SEC can be learned by observing human manipulations, transferred to a new scenario, and then reproduced by the machine. © SAGE Publications 2011.The research leading to these results has received funding from the European Community's Seventh Framework Programme FP7=2007-2013 - Challenge 2 - Cognitive Systems, Interaction, Robotics - under grant agreement No 247947 - GARNICS. B.D. acknowledges support from the Spanish Ministry for Science and Innovation via a Ramon y Cajal Fellowship.Peer ReviewedSage Publications2014201420112014info:eu-repo/semantics/articlehttp://purl.org/coar/resource_type/c_6501Postprintinfo:eu-repo/semantics/acceptedVersionhttp://hdl.handle.net/10261/96813reponame:DIGITAL.CSIC. Repositorio Institucional del CSICinstname:Consejo Superior de Investigaciones Científicas (CSIC)Inglés#PLACEHOLDER_PARENT_METADATA_VALUE#info:eu-repo/grantAgreement/EC/FP7/247947http://dx.doi.org/10.1177/0278364911410459info:eu-repo/semantics/openAccessoai:digital.csic.es:10261/968132026-05-22T06:33:51Z |
| dc.title.none.fl_str_mv |
Learning the semantics of object-action relations by observation |
| title |
Learning the semantics of object-action relations by observation |
| spellingShingle |
Learning the semantics of object-action relations by observation Aksoy, Eren Erdal Action recognition Affordances Object categorization Semantic scene graphs Unsupervised learning Object–action complexes (OACs) |
| title_short |
Learning the semantics of object-action relations by observation |
| title_full |
Learning the semantics of object-action relations by observation |
| title_fullStr |
Learning the semantics of object-action relations by observation |
| title_full_unstemmed |
Learning the semantics of object-action relations by observation |
| title_sort |
Learning the semantics of object-action relations by observation |
| dc.creator.none.fl_str_mv |
Aksoy, Eren Erdal Abramov, Alexey Dörr, Johannes Ning, Kejun Dellen, Babette Wörgötter, Florentin |
| author |
Aksoy, Eren Erdal |
| author_facet |
Aksoy, Eren Erdal Abramov, Alexey Dörr, Johannes Ning, Kejun Dellen, Babette Wörgötter, Florentin |
| author_role |
author |
| author2 |
Abramov, Alexey Dörr, Johannes Ning, Kejun Dellen, Babette Wörgötter, Florentin |
| author2_role |
author author author author author |
| dc.subject.none.fl_str_mv |
Action recognition Affordances Object categorization Semantic scene graphs Unsupervised learning Object–action complexes (OACs) |
| topic |
Action recognition Affordances Object categorization Semantic scene graphs Unsupervised learning Object–action complexes (OACs) |
| description |
Recognizing manipulations performed by a human and the transfer and execution of this by a robot is a difficult problem. We address this in the current study by introducing a novel representation of the relations between objects at decisive time points during a manipulation. Thereby, we encode the essential changes in a visual scenery in a condensed way such that a robot can recognize and learn a manipulation without prior object knowledge. To achieve this we continuously track image segments in the video and construct a dynamic graph sequence. Topological transitions of those graphs occur whenever a spatial relation between some segments has changed in a discontinuous way and these moments are stored in a transition matrix called the semantic event chain (SEC). We demonstrate that these time points are highly descriptive for distinguishing between different manipulations. Employing simple sub-string search algorithms, SECs can be compared and type-similar manipulations can be recognized with high confidence. As the approach is generic, statistical learning can be used to find the archetypal SEC of a given manipulation class. The performance of the algorithm is demonstrated on a set of real videos showing hands manipulating various objects and performing different actions. In experiments with a robotic arm, we show that the SEC can be learned by observing human manipulations, transferred to a new scenario, and then reproduced by the machine. © SAGE Publications 2011. |
| publishDate |
2011 |
| dc.date.none.fl_str_mv |
2011 2014 2014 2014 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article http://purl.org/coar/resource_type/c_6501 Postprint info:eu-repo/semantics/acceptedVersion |
| format |
article |
| status_str |
acceptedVersion |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10261/96813 |
| url |
http://hdl.handle.net/10261/96813 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
#PLACEHOLDER_PARENT_METADATA_VALUE# info:eu-repo/grantAgreement/EC/FP7/247947 http://dx.doi.org/10.1177/0278364911410459 |
| dc.rights.none.fl_str_mv |
info:eu-repo/semantics/openAccess |
| eu_rights_str_mv |
openAccess |
| dc.publisher.none.fl_str_mv |
Sage Publications |
| publisher.none.fl_str_mv |
Sage Publications |
| dc.source.none.fl_str_mv |
reponame:DIGITAL.CSIC. Repositorio Institucional del CSIC instname:Consejo Superior de Investigaciones Científicas (CSIC) |
| instname_str |
Consejo Superior de Investigaciones Científicas (CSIC) |
| reponame_str |
DIGITAL.CSIC. Repositorio Institucional del CSIC |
| collection |
DIGITAL.CSIC. Repositorio Institucional del CSIC |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869416033237860352 |
| score |
15,812455 |