Exploring Graph Neural Networks for Video Action Segmentation

A common challenge in computer vision is the applicability of algorithms developed in a controlled dataset to real-world problems, such as unscripted or uncontrolled videos. Graph neural networks (GNNs) have emerged as a promising tool to solve this kind of problem. This master's thesis explore...

Descripción completa

Detalles Bibliográficos
Autor: Vaccher Gomez, Jordi
Tipo de recurso: tesis de maestría
Fecha de publicación:2023
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/399778
Acceso en línea:https://hdl.handle.net/2117/399778
Access Level:acceso abierto
Palabra clave:Machine learning
Neural networks (Computer science)
Deep learning (Machine learning)
Graph Neural Networks
machine learning
deep learning
action segmentation
video data
Aprenentatge automàtic
Xarxes neuronals (Informàtica)
Aprenentatge profund
Àrees temàtiques de la UPC::Informàtica::Intel·ligència artificial
Descripción
Sumario:A common challenge in computer vision is the applicability of algorithms developed in a controlled dataset to real-world problems, such as unscripted or uncontrolled videos. Graph neural networks (GNNs) have emerged as a promising tool to solve this kind of problem. This master's thesis explores the use of graph neural networks for the classification of actions and activities in video data. Specifically, the study focuses on the Breakfast dataset, which features two types of labels: activity labels for each video and action labels for each frame. In this work, we investigate various methods for action and activity classification and propose a novel architecture that employs a graph neural network for unsupervised embedding extraction trained to separate the intra-video action classes, improving the baseline performance. The proposed GNN-based embedding extractor architecture shows the capability of graph-based techniques to improve our comprehension of complex events and actions in video data.