arXiv Open Access 2025

Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction

Wenxuan Wu Shuai Wang Xixin Wu Helen Meng Haizhou Li

Lihat Sumber

Abstrak

Audio-visual target speaker extraction (AV-TSE) models primarily rely on target visual cues to isolate the target speaker's voice from others. We know that humans leverage linguistic knowledge, such as syntax and semantics, to support speech perception. Inspired by this, we explore the potential of pre-trained speech-language models (PSLMs) and pre-trained language models (PLMs) as auxiliary knowledge sources for AV-TSE. In this study, we propose incorporating the linguistic constraints from PSLMs or PLMs for the AV-TSE model as additional supervision signals. Without introducing any extra computational cost during inference, the proposed approach consistently improves speech quality and intelligibility. Furthermore, we evaluate our method in multi-language settings and visual cue-impaired scenarios and show robust performance gains.

Topik & Kata Kunci

cs.SD cs.LG cs.MM eess.AS

Penulis (5)

Wenxuan Wu

Shuai Wang

Xixin Wu

Helen Meng

Haizhou Li

Format Sitasi

APA MLA BibTeX

Wu, W., Wang, S., Wu, X., Meng, H., Li, H. (2025). Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction. https://arxiv.org/abs/2506.09792

Akses Cepat

Lihat di Sumber

Informasi Jurnal

Tahun Terbit: 2025
Bahasa: en
Sumber Database: arXiv
Akses: Open Access ✓