Efficient training of large neural networks for language modeling

Эффективное обучение крупных нейронных сетей для языкового моделирования
Holger Schwenk
2005-02-28

DARPA Rich Transcriptions / conversational speech recognizerefficient traininglarge vocabulary speech recognitionlattice rescoringneural network language model
Recently there has been increasing interest in using neural networks for language modeling. In contrast to the well-known backoff n-gram language models, the neural network approach tries to limit the data sparseness problem by performing the estimation in a continuous space, allowing by this means smooth interpolations. The complexity to train such a model and to calculate one n-gram probability is however several orders of magnitude higher than for the backoff models, making the new approach difficult to use in real applications. In this paper several techniques are presented that allow the use of a neural network language model in a large vocabulary speech recognition system, in particular very, fast lattice rescoring and efficient training of large neural networks on training corpora of over 10 million words. The described approach achieves significant word error reductions with respect to a carefully tuned 4-gram backoff language model in a state of the art conversational speech recognizer for the DARPA rich transcriptions evaluations.
1
Neural network language models estimate probabilities in a continuous space, reducing data sparsity via smooth interpolations compared to backoff n-gram models.
2
The paper introduces techniques enabling efficient training of large neural networks on corpora over 10 million words.
3
The paper presents very fast lattice rescoring suitable for large-vocabulary speech recognition using neural network language models.
4
Training and computing n-gram probabilities with neural networks is several orders of magnitude more computationally expensive than backoff models.
5
Using the described approach yields significant word error rate reductions versus a carefully tuned 4-gram backoff model in a state-of-the-art conversational speech recognizer for DARPA RT evaluations.

Large neural network language models trained on corpora of over 10 million words

Efficient training methods and fast lattice-rescoring techniques to enable use of large neural network language models in large-vocabulary conversational speech recognition and reduce word error rate compared to 4-gram backoff models

Publication Details
Publication Date
2005-02-28
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Holger Schwenk
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%