arXiv Open Access 2026

SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality

Qinghao Guan Yuchen Pan Donghao Li Zishi Zhang Yiyang Chen +4 lainnya

Lihat Sumber

Abstrak

In religion and theology studies, spirituality has garnered significant research attention for the reason that it not only transcends culture but offers unique experience to each individual. However, social scientists often rely on limited datasets, which are basically unavailable online. In this study, we collaborated with social scientists to develop a high-quality multimedia multi-modal datasets, \textbf{SACRED}, in which the faithfulness of classification is guaranteed. Using \textbf{SACRED}, we evaluated the performance of 13 popular LLMs as well as traditional rule-based and fine-tuned approaches. The result suggests DeepSeek-V3 model performs well in classifying such abstract concepts (i.e., 79.19\% accuracy in the Quora test set), and the GPT-4o-mini model surpassed the other models in the vision tasks (63.99\% F1 score). Purportedly, this is the first annotated multi-modal dataset from online spirituality communication. Our study also found a new type of connectedness which is valuable for communication science studies.

Topik & Kata Kunci

cs.CL cs.MM

Penulis (9)

Qinghao Guan

Yuchen Pan

Donghao Li

Zishi Zhang

Yiyang Chen

Lu Li

Flaminia Canu

Emilia Volkart

Gerold Schneider

Format Sitasi

APA MLA BibTeX

Guan, Q., Pan, Y., Li, D., Zhang, Z., Chen, Y., Li, L. et al. (2026). SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality. https://arxiv.org/abs/2603.27331

Akses Cepat

Lihat di Sumber

Informasi Jurnal

Tahun Terbit: 2026
Bahasa: en
Sumber Database: arXiv
Akses: Open Access ✓