Parallel Corpora in Focus: An Account of Current Achievements and Challenges

This article provides an overview of the development of parallel corpora over the past three decades. The corpus-based approach has transformed linguistic research since the 1990s, particularly with the emergence of parallel corpora, which have proven indispensable in both linguistic and computation...

ver descrição completa

Detalhes bibliográficos
Autores: Sánchez Nieto, María Teresa, Doval Reixa, Irene
Formato: capítulo de livro
Fecha de publicación:2019
País:España
Recursos:Universidad de Santiago de Compostela (USC)
Repositorio:Minerva. Repositorio Institucional de la Universidad de Santiago de Compostela
Idioma:inglés
OAI Identifier:oai:minerva.usc.gal:10347/39209
Acesso em linha:https://hdl.handle.net/10347/39209
Access Level:acceso abierto
Palavra-chave:Corpus linguistics
Parallel corpora
Applications of parallel corpora
5701 Lingüística aplicada
Descrição
Resumo:This article provides an overview of the development of parallel corpora over the past three decades. The corpus-based approach has transformed linguistic research since the 1990s, particularly with the emergence of parallel corpora, which have proven indispensable in both linguistic and computational fields. Early efforts, like the Yugoslav Serbo-Croatian–English Contrastive Project and the Canadian Hansard Corpus, laid the groundwork for modern parallel corpora, exemplified by the English-Norwegian Parallel Corpus (ENPC) and its Swedish counterpart. Over time, major initiatives, such as the Europarl Corpus and OPUS, have significantly expanded multilingual resources for research and applications. The paper describes the main milestones in their evolution, highlighting pioneering works. Additionally, parallel corpora have become an indispensable resource for a wide range of applications. The core focus of the article lies in exploring these applications. Five major fields of application are identified, each with its own specific needs and user groups: (a) basic research in contrastive linguistics and translation studies, (b) lexicography, (c) translation practice, (d) foreign language and translation teaching, and (e) natural language processing, particularly machine translation. The processing and development of parallel corpora remain complex but essential. Researchers balance the need for technological expertise with collaboration or specialized tools for compiling, annotating, and aligning corpora. Innovations in alignment, from sentence to word level, enhance usability, supporting advanced applications like semantic annotation and dependency parsing. Reusability is also a priority, with many projects adapting existing corpora for new purposes. Tools for indexing and querying continue to evolve, incorporating capabilities beyond standard concordancing to include collocation analysis and statistical exploration. This dynamic field reflects the synergy between linguistic and computational traditions, with parallel corpora playing a pivotal role in advancing cross-linguistic research, translation studies, and language education. The ongoing development and reuse of these resources will further enhance their impact across disciplines.