An approach for discovering keywords from Spanish tweets using Wikipedia

Most approaches to keywords discovery when analyzing microblogging messages (among them those from Twitter) are based on statistical and lexical information about the words that compose the text. The lack of context in the short messages can be problematic due to the low co-occurrence of words. In t...

Full description

Bibliographic Details
Authors: Ayala Hernández, Daniel, Roldán Salvador, Juan Carlos, Ruiz Cortés, David, Ortega Gallego, Fernando
Format: article
Status:Published version
Publication Date:2015
Country:España
Institution:Universidad de Sevilla (US)
Repository:idUS. Depósito de Investigación de la Universidad de Sevilla
OAI Identifier:oai:idus.us.es:11441/49141
Online Access:http://hdl.handle.net/11441/49141
https://doi.org/10.14201/ADCAIJ2015427388
Access Level:Open access
Keyword:Twitter
Social Media Analysis
Wikipedia
Keywords Discovery
Description
Summary:Most approaches to keywords discovery when analyzing microblogging messages (among them those from Twitter) are based on statistical and lexical information about the words that compose the text. The lack of context in the short messages can be problematic due to the low co-occurrence of words. In this paper, we present a new approach for keywords discovering from Spanish tweets based on the addition of context information using Wikipedia as a knowledge base. We present four different ways to use Wikipedia and two ways to rank the new keywords. We have tested these strategies using more than 60000 Spanish tweets, measuring performance and analyzing particularities of each strategy.