Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

Уточнение псевдоаудио-подсказок с выравниванием речи и текста для адаптации по домену только по тексту в ASR на основе LLM
Ryo Magoshi, Takashi Maekaku, Yusuke Shinohara
2026-05-14

LLM-based ASRout-of-vocabulary coveragepseudo-audio promptsspeech-text alignmenttext-only domain adaptation
LLM-based automatic speech recognition models demonstrate strong performance by connecting audio encoders and LLMs. However, data scarcity of paired speech and transcription often hinders their adaptation to new domains, making text-only domain adaptation crucial. Existing methods typically rely on either fine-tuning the LLM alone or employing pseudo-audio prompts. The former neglects essential acoustic context, while the latter either suffers from limited scalability in data-scarce conditions, or yields inexpressive prompts by leveraging only textual features, ignoring audio modality. To address this, we propose an enhanced framework that explicitly models speech-text alignment. Our method efficiently generates highly expressive pseudo-audio prompts that bridges the modality gap, enabling effective target-domain adaptation. Experiments demonstrate that our approach outperforms existing text-only methods, improving both overall error rates and out-of-vocabulary coverage.
1
Existing text-only domain adaptation methods for LLM-based ASR either fine-tune only the LLM (losing acoustic context) or use pseudo-audio prompts that are either not scalable or inexpressive.
2
Experiments show the proposed approach outperforms existing text-only methods by reducing overall error rates and improving out-of-vocabulary coverage.
3
The authors propose a framework that explicitly models speech-text alignment to generate more expressive pseudo-audio prompts bridging the audio-text modality gap.
4
Their method enables efficient generation of highly expressive pseudo-audio prompts, improving target-domain adaptation without requiring paired speech-transcription data.

LLM-based automatic speech recognition system adapted to a target domain using text-only data via pseudo-audio prompts

Generation and refinement of expressive pseudo-audio prompts via explicit speech–text alignment to enable effective text-only domain adaptation and improve ASR error rates and out-of-vocabulary coverage

Publication Details
Publication Date
2026-05-14
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Ryo Magoshi
Takashi Maekaku
Yusuke Shinohara
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%