Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR
Уточнение псевдоаудио-подсказок с выравниванием речи и текста для адаптации по домену только по тексту в ASR на основе LLM
2026-05-14
SCID: 54.1/9mggpu3z
Discuss with AI
LLM-based ASRout-of-vocabulary coveragepseudo-audio promptsspeech-text alignmenttext-only domain adaptation
Figures from the paper
Abstract (AI)
LLM-based automatic speech recognition models demonstrate strong performance by connecting audio encoders and LLMs. However, data scarcity of paired speech and transcription often hinders their adaptation to new domains, making text-only domain adaptation crucial. Existing methods typically rely on either fine-tuning the LLM alone or employing pseudo-audio prompts. The former neglects essential acoustic context, while the latter either suffers from limited scalability in data-scarce conditions, or yields inexpressive prompts by leveraging only textual features, ignoring audio modality. To address this, we propose an enhanced framework that explicitly models speech-text alignment. Our method efficiently generates highly expressive pseudo-audio prompts that bridges the modality gap, enabling effective target-domain adaptation. Experiments demonstrate that our approach outperforms existing text-only methods, improving both overall error rates and out-of-vocabulary coverage.
Key Findings
1
Existing text-only domain adaptation methods for LLM-based ASR either fine-tune only the LLM (losing acoustic context) or use pseudo-audio prompts that are either not scalable or inexpressive.
2
Experiments show the proposed approach outperforms existing text-only methods by reducing overall error rates and improving out-of-vocabulary coverage.
3
The authors propose a framework that explicitly models speech-text alignment to generate more expressive pseudo-audio prompts bridging the audio-text modality gap.
4
Their method enables efficient generation of highly expressive pseudo-audio prompts, improving target-domain adaptation without requiring paired speech-transcription data.
Research Object
LLM-based automatic speech recognition system adapted to a target domain using text-only data via pseudo-audio prompts
Research Subject
Generation and refinement of expressive pseudo-audio prompts via explicit speech–text alignment to enable effective text-only domain adaptation and improve ASR error rates and out-of-vocabulary coverage
Publication Details
Publication Date
2026-05-14
Journal
Publisher
ISSN
Cited by
0
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest