arXiv Open Access 2025

Cleaning English Abstracts of Scientific Publications

Michael E. Rose Nils A. Herrmann Sebastian Erhardt
Lihat Sumber

Abstrak

Scientific abstracts are often used as proxies for the content and thematic focus of research publications. However, a significant share of published abstracts contains extraneous information-such as publisher copyright statements, section headings, author notes, registrations, and bibliometric or bibliographic metadata-that can distort downstream analyses, particularly those involving document similarity or textual embeddings. We introduce an open-source, easy-to-integrate language model designed to clean English-language scientific abstracts by automatically identifying and removing such clutter. We demonstrate that our model is both conservative and precise, alters similarity rankings of cleaned abstracts and improves information content of standard-length embeddings.

Topik & Kata Kunci

Penulis (3)

M

Michael E. Rose

N

Nils A. Herrmann

S

Sebastian Erhardt

Format Sitasi

Rose, M.E., Herrmann, N.A., Erhardt, S. (2025). Cleaning English Abstracts of Scientific Publications. https://arxiv.org/abs/2512.24459

Akses Cepat

Lihat di Sumber
Informasi Jurnal
Tahun Terbit
2025
Bahasa
en
Sumber Database
arXiv
Akses
Open Access ✓