Squeezing Bottlenecks: Exploring the Limits of Autoencoder Semantic Representation Capabilities

[EN] We present a comprehensive study on the use of autoencoders for modelling text data, in which (differently from previous studies) we focus our attention on the various issues. We explore the suitability of two different models binary deep autencoders (bDA) and replicated-softmax deep autencoder...

Descripción completa

Detalles Bibliográficos
Autores: Gupta, Parth Alokkumar, Banchs, Rafael, Rosso, Paolo
Tipo de recurso: artículo
Fecha de publicación:2016
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/82646
Acceso en línea:https://riunet.upv.es/handle/10251/82646
Access Level:acceso abierto
Palabra clave:Text representation
Deep autoencoder
LENGUAJES Y SISTEMAS INFORMATICOS
Descripción
Sumario:[EN] We present a comprehensive study on the use of autoencoders for modelling text data, in which (differently from previous studies) we focus our attention on the various issues. We explore the suitability of two different models binary deep autencoders (bDA) and replicated-softmax deep autencoders (rsDA) for constructing deep autoencoders for text data at the sentence level. We propose and evaluate two novel metrics for better assessing the text-reconstruction capabilities of autoencoders. We propose an automatic method to find the critical bottleneck dimensionality for text representations (below which structural information is lost); and finally we conduct a comparative evaluation across different languages, exploring the regions of critical bottleneck dimensionality and its relationship to language perplexity. & 2015 Elsevier B.V. All rights reserved.