Strategies for training large scale neural network language models
Стратегии обучения крупномасштабных языковых моделей на основе нейронных сетей
2011-12-01
SCID: 54.1/y2wfqv34
Discuss with AI
4-gram baseline trained on 400M tokenshash-based maximum entropy modelintegrated maxent within neural networklarge scale neural network language modelstraining data sorting by relevanceword error rate reduction on English Broadcast News
Figures from the paper
Abstract (AI)
We describe how to effectively train neural network based language models on large data sets. Fast convergence during training and better overall performance is observed when the training data are sorted by their relevance. We introduce hash-based implementation of a maximum entropy model, that can be trained as a part of the neural network model. This leads to significant reduction of computational complexity. We achieved around 10% relative reduction of word error rate on English Broadcast News speech recognition task, against large 4-gram model trained on 400M tokens.
Key Findings
1
A hash-based implementation of a maximum entropy component can be integrated into the neural network language model.
2
Integrating the hash-based maximum entropy model substantially reduces computational complexity.
3
On English Broadcast News speech recognition, the proposed approach achieved about a 10% relative reduction in word error rate compared to a large 4-gram model trained on 400M tokens.
4
Sorting training data by relevance yields faster convergence and better overall performance for neural network language model training.
Research Object
Neural network–based language models trained on large datasets
Research Subject
Training strategies and modifications (data sorting by relevance and hash-based max-entropy component) that improve convergence, reduce computational complexity, and lower word error rate
Publication Details
Publication Date
2011-12-01
Journal
Publisher
ISSN
Access Type
Author Information
Download PDF
Subscribe to digest