arXiv Open Access 2024

Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Modan Tailleur Junwon Lee Mathieu Lagrange Keunwoo Choi Laurie M. Heller +2 lainnya
Lihat Sumber

Abstrak

This paper explores whether considering alternative domain-specific embeddings to calculate the Fréchet Audio Distance (FAD) metric can help the FAD to correlate better with perceptual ratings of environmental sounds. We used embeddings from VGGish, PANNs, MS-CLAP, L-CLAP, and MERT, which are tailored for either music or environmental sound evaluation. The FAD scores were calculated for sounds from the DCASE 2023 Task 7 dataset. Using perceptual data from the same task, we find that PANNs-WGM-LogMel produces the best correlation between FAD scores and perceptual ratings of both audio quality and perceived fit with a Spearman correlation higher than 0.5. We also find that music-specific embeddings resulted in significantly lower results. Interestingly, VGGish, the embedding used for the original Fréchet calculation, yielded a correlation below 0.1. These results underscore the critical importance of the choice of embedding for the FAD metric design.

Topik & Kata Kunci

Penulis (7)

M

Modan Tailleur

J

Junwon Lee

M

Mathieu Lagrange

K

Keunwoo Choi

L

Laurie M. Heller

K

Keisuke Imoto

Y

Yuki Okamoto

Format Sitasi

Tailleur, M., Lee, J., Lagrange, M., Choi, K., Heller, L.M., Imoto, K. et al. (2024). Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant. https://arxiv.org/abs/2403.17508

Akses Cepat

Lihat di Sumber
Informasi Jurnal
Tahun Terbit
2024
Bahasa
en
Sumber Database
arXiv
Akses
Open Access ✓