DOAJ Open Access 2025

BMWP: the first Bengali math word problems dataset for operation prediction and solving

Sanchita Mondal Debnarayan Khatua Sourav Mandal Dilip K. Prasad Arif Ahmed Sekh

Abstrak

Abstract Solving math word problems of varying complexities is one of the most challenging and exciting research questions in artificial intelligence (AI), particularly in natural language processing (NLP) and machine learning (ML). Foundational language models such as GPT must be evaluated for intelligence, and solving word problems is a key method for this assessment. These problems become especially difficult when presented in low-resource regional languages such as Bengali. Word problem solving integrates the cognitive domains of language processing, comprehension, and transformation into real-world solutions. During the past decade, advances in AI and machine learning have significantly progressed in addressing this complex issue. Although researchers worldwide have primarily utilized datasets in English and some in Chinese, there has been a lack of standard datasets for low-resource languages such as Bengali. In this pioneering study, we introduce the first Bengali Math Word Problem Benchmark Data Set (BMWP), comprising 8653 word problems. We detail the creation of this dataset and the benchmarking methods employed. Furthermore, we investigate operation prediction from Bengali word problems using state-of-the-art deep learning (DL) techniques. We implemented and compared various standard DL-based neural network architectures, achieving an accuracy of $$92 \pm 2\%$$ 92 ± 2 % . The data set and the code will be available at https://github.com/SanchitaMondal/BMWP .

Topik & Kata Kunci

Computational linguistics. Natural language processing Electronic computers. Computer science

Penulis (5)

Sanchita Mondal

Debnarayan Khatua

Sourav Mandal

Dilip K. Prasad

Arif Ahmed Sekh

Format Sitasi

APA MLA BibTeX

Mondal, S., Khatua, D., Mandal, S., Prasad, D.K., Sekh, A.A. (2025). BMWP: the first Bengali math word problems dataset for operation prediction and solving. https://doi.org/10.1007/s44163-025-00243-7

Akses Cepat

PDF tidak tersedia langsung

Cek di sumber asli →

Lihat di Sumber doi.org/10.1007/s44163-025-00243-7

Informasi Jurnal

Tahun Terbit: 2025
Sumber Database: DOAJ
DOI: 10.1007/s44163-025-00243-7
Akses: Open Access ✓