Efficient training of large neural networks for language modeling
Эффективное обучение крупных нейронных сетей для языкового моделирования
2005-02-28
SCID: 54.1/4fujxh38
Discuss with AI
DARPA Rich Transcriptions / conversational speech recognizerefficient traininglarge vocabulary speech recognitionlattice rescoringneural network language model
Figures from the paper
Abstract (AI)
Recently there has been increasing interest in using neural networks for language modeling. In contrast to the well-known backoff n-gram language models, the neural network approach tries to limit the data sparseness problem by performing the estimation in a continuous space, allowing by this means smooth interpolations. The complexity to train such a model and to calculate one n-gram probability is however several orders of magnitude higher than for the backoff models, making the new approach difficult to use in real applications. In this paper several techniques are presented that allow the use of a neural network language model in a large vocabulary speech recognition system, in particular very, fast lattice rescoring and efficient training of large neural networks on training corpora of over 10 million words. The described approach achieves significant word error reductions with respect to a carefully tuned 4-gram backoff language model in a state of the art conversational speech recognizer for the DARPA rich transcriptions evaluations.
Key Findings
1
Neural network language models estimate probabilities in a continuous space, reducing data sparsity via smooth interpolations compared to backoff n-gram models.
2
The paper introduces techniques enabling efficient training of large neural networks on corpora over 10 million words.
3
The paper presents very fast lattice rescoring suitable for large-vocabulary speech recognition using neural network language models.
4
Training and computing n-gram probabilities with neural networks is several orders of magnitude more computationally expensive than backoff models.
5
Using the described approach yields significant word error rate reductions versus a carefully tuned 4-gram backoff model in a state-of-the-art conversational speech recognizer for DARPA RT evaluations.
Research Object
Large neural network language models trained on corpora of over 10 million words
Research Subject
Efficient training methods and fast lattice-rescoring techniques to enable use of large neural network language models in large-vocabulary conversational speech recognition and reduce word error rate compared to 4-gram backoff models
Publication Details
Publication Date
2005-02-28
Journal
Publisher
ISSN
Access Type
Author Information
Download PDF
Subscribe to digest