Kaldi Speech Recognition Toolkit
Инструментарий распознавания речи Kaldi
2024-01-01
SCID: 54.1/rjmc9cwq
Discuss with AI
Kaldi Speech Recognition ToolkitOpenFstacoustic modelingfinite-state transducerssubspace Gaussian mixture models
Figures from the paper
Abstract (AI)
Abstract—We describe the design of Kaldi, a free, open-source toolkit for speech recognition research. Kaldi provides a speech recognition system based on finite-state transducers (using the freely available OpenFst), together with detailed documentation and scripts for building complete recognition systems. Kaldi is written is C++, and the core library supports modeling of arbitrary phonetic-context sizes, acoustic modeling with subspace Gaussian mixture models (SGMM) as well as standard Gaussian mixture models, together with all commonly used linear and affine transforms. Kaldi is released under the Apache License v2.0, which is highly nonrestrictive, making it suitable for a wide community of users. I.
Key Findings
1
Its C++ core supports arbitrary phonetic-context sizes, SGMM and standard GMM acoustic models, and commonly used linear and affine transforms.
2
Kaldi includes documentation and scripts for constructing complete speech recognition systems, supporting reproducible system development.
3
Kaldi is a free, open-source speech recognition toolkit designed specifically for speech recognition research.
4
Kaldi is distributed under the permissive Apache License 2.0, enabling broad use by the speech recognition community.
5
The toolkit implements recognition systems using finite-state transducers through the freely available OpenFst library.
Research Object
Kaldi, a free and open-source speech recognition toolkit
Research Subject
The toolkit’s design and capabilities for building complete speech recognition systems, including finite-state-transducer-based recognition and acoustic modeling
Publication Details
Publication Date
2024-01-01
Journal
Publisher
ISSN
Cited by
4890
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest
Cited by7
X-Vectors: Robust DNN Embeddings for Speaker Recognition2018
A study on data augmentation of reverberant speech for robust speech recognition2017
TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech2021
Analysis of Length Normalization in End-to-End Speaker Verification System2018
Barlow Twins self-supervised learning for robust speaker recognition2022
Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems2020
Deep neural network-based speaker embeddings for end-to-end speaker verification2016