Stance and engagement in English and Spanish journalistic texts : towards a reliable annotation scheme for linguistic and computational purposes

The paper describes the process of validating a reliable annotation scheme for the categories of Stance and Engagement in English and Spanish using a bilingual sample of English-Spanish journalistic texts extracted from the MULTINOT corpus (Lavid et al. 2015). The bilingual sample includes three dif...

Full description

Bibliographic Details
Authors: Moratón Gutiérrez, Lara, Lavid López, María Julia
Format: article
Publication Date:2018
Country:España
Institution:Universidad Complutense de Madrid (UCM)
Repository:Docta Complutense
Language:English
OAI Identifier:oai:docta.ucm.es:20.500.14352/96472
Online Access:https://hdl.handle.net/20.500.14352/96472
Access Level:Open access
Keyword:Filología inglesa
5701.04 Lingüística Informatizada
Description
Summary:The paper describes the process of validating a reliable annotation scheme for the categories of Stance and Engagement in English and Spanish using a bilingual sample of English-Spanish journalistic texts extracted from the MULTINOT corpus (Lavid et al. 2015). The bilingual sample includes three different newspaper genres: news reports, editorials and letters to the editor. Following the generic annotation pipeline proposed by Hovy and Lavid (2010), the paper describes the different steps to validate an annotation scheme to capture the main features of Stance and Engagement and their realizations in English and Spanish. This includes the instantiation of the theoretical categories, the development of annotation guidelines and the performance of an interannotation agreement study on a small training corpus to measure the reliability of the proposed tags. The paper also describes the results of the annotation of a larger corpus using the validated scheme (Moratón 2015). This reveals interesting patterns of variation in the distribution of Stance and Engagement in the three newspaper genres, which can be fruitfully used for contrastive linguistic and computational purposes.