Exploring Graph Neural Networks for Video Action Segmentation
A common challenge in computer vision is the applicability of algorithms developed in a controlled dataset to real-world problems, such as unscripted or uncontrolled videos. Graph neural networks (GNNs) have emerged as a promising tool to solve this kind of problem. This master's thesis explore...
| Autor: | |
|---|---|
| Tipo de recurso: | tesis de maestría |
| Fecha de publicación: | 2023 |
| País: | España |
| Institución: | Universitat Politècnica de Catalunya (UPC) |
| Repositorio: | UPCommons. Portal del coneixement obert de la UPC |
| Idioma: | inglés |
| OAI Identifier: | oai:upcommons.upc.edu:2117/399778 |
| Acceso en línea: | https://hdl.handle.net/2117/399778 |
| Access Level: | acceso abierto |
| Palabra clave: | Machine learning Neural networks (Computer science) Deep learning (Machine learning) Graph Neural Networks machine learning deep learning action segmentation video data Aprenentatge automàtic Xarxes neuronals (Informàtica) Aprenentatge profund Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial |
| Sumario: | A common challenge in computer vision is the applicability of algorithms developed in a controlled dataset to real-world problems, such as unscripted or uncontrolled videos. Graph neural networks (GNNs) have emerged as a promising tool to solve this kind of problem. This master's thesis explores the use of graph neural networks for the classification of actions and activities in video data. Specifically, the study focuses on the Breakfast dataset, which features two types of labels: activity labels for each video and action labels for each frame. In this work, we investigate various methods for action and activity classification and propose a novel architecture that employs a graph neural network for unsupervised embedding extraction trained to separate the intra-video action classes, improving the baseline performance. The proposed GNN-based embedding extractor architecture shows the capability of graph-based techniques to improve our comprehension of complex events and actions in video data. |
|---|