arXiv Open Access 2023

Environmental sound synthesis from vocal imitations and sound event labels

Yuki Okamoto Keisuke Imoto Shinnosuke Takamichi Ryotaro Nagase Takahiro Fukumori +1 lainnya

Lihat Sumber

Abstrak

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as rhythm and pitch, using vocal imitations, which cannot be expressed by conventional input information, such as sound event labels, images, or texts, in an environmental sound synthesis model. In this paper, we propose a framework for environmental sound synthesis from vocal imitations and sound event labels based on a framework of a vector quantized encoder and the Tacotron2 decoder. Using vocal imitations is expected to control the pitch and rhythm of the synthesized sound, which only sound event labels cannot control. Our objective and subjective experimental results show that vocal imitations effectively control the pitch and rhythm of synthesized sounds.

Topik & Kata Kunci

cs.SD eess.AS

Penulis (6)

Yuki Okamoto

Keisuke Imoto

Shinnosuke Takamichi

Ryotaro Nagase

Takahiro Fukumori

Yoichi Yamashita

Format Sitasi

APA MLA BibTeX

Okamoto, Y., Imoto, K., Takamichi, S., Nagase, R., Fukumori, T., Yamashita, Y. (2023). Environmental sound synthesis from vocal imitations and sound event labels. https://arxiv.org/abs/2305.00302

Akses Cepat

Lihat di Sumber

Informasi Jurnal

Tahun Terbit: 2023
Bahasa: en
Sumber Database: arXiv
Akses: Open Access ✓