arXiv Open Access 2025

Towards Benchmarking Foundation Models for Tabular Data With Text

Martin Mráz Breenda Das Anshul Gupta Lennart Purucker Frank Hutter

Lihat Sumber

Abstrak

Foundation models for tabular data are rapidly evolving, with increasing interest in extending them to support additional modalities such as free-text features. However, existing benchmarks for tabular data rarely include textual columns, and identifying real-world tabular datasets with semantically rich text features is non-trivial. We propose a series of simple yet effective ablation-style strategies for incorporating text into conventional tabular pipelines. Moreover, we benchmark how state-of-the-art tabular foundation models can handle textual data by manually curating a collection of real-world tabular datasets with meaningful textual features. Our study is an important step towards improving benchmarking of foundation models for tabular data with text.

Topik & Kata Kunci

cs.LG

Penulis (5)

Martin Mráz

Breenda Das

Anshul Gupta

Lennart Purucker

Frank Hutter

Format Sitasi

APA MLA BibTeX

Mráz, M., Das, B., Gupta, A., Purucker, L., Hutter, F. (2025). Towards Benchmarking Foundation Models for Tabular Data With Text. https://arxiv.org/abs/2507.07829

Akses Cepat

Lihat di Sumber

Informasi Jurnal

Tahun Terbit: 2025
Bahasa: en
Sumber Database: arXiv
Akses: Open Access ✓