arXiv Open Access 2024

EUFCC-340K: A Faceted Hierarchical Dataset for Metadata Annotation in GLAM Collections

Francesc Net Marc Folia Pep Casals Andrew D. Bagdanov Lluis Gomez
Lihat Sumber

Abstrak

In this paper, we address the challenges of automatic metadata annotation in the domain of Galleries, Libraries, Archives, and Museums (GLAMs) by introducing a novel dataset, EUFCC340K, collected from the Europeana portal. Comprising over 340,000 images, the EUFCC340K dataset is organized across multiple facets: Materials, Object Types, Disciplines, and Subjects, following a hierarchical structure based on the Art & Architecture Thesaurus (AAT). We developed several baseline models, incorporating multiple heads on a ConvNeXT backbone for multi-label image tagging on these facets, and fine-tuning a CLIP model with our image text pairs. Our experiments to evaluate model robustness and generalization capabilities in two different test scenarios demonstrate the utility of the dataset in improving multi-label classification tools that have the potential to alleviate cataloging tasks in the cultural heritage sector.

Topik & Kata Kunci

Penulis (5)

F

Francesc Net

M

Marc Folia

P

Pep Casals

A

Andrew D. Bagdanov

L

Lluis Gomez

Format Sitasi

Net, F., Folia, M., Casals, P., Bagdanov, A.D., Gomez, L. (2024). EUFCC-340K: A Faceted Hierarchical Dataset for Metadata Annotation in GLAM Collections. https://arxiv.org/abs/2406.02380

Akses Cepat

Lihat di Sumber
Informasi Jurnal
Tahun Terbit
2024
Bahasa
en
Sumber Database
arXiv
Akses
Open Access ✓