Semantic Scholar Open Access 2022 16 sitasi

Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono

A. Akther Md. Shymon Islam Hafsa Sultana A. K. Z. Rasel Rahman Sujan Kumar Saha +2 lainnya

Lihat Sumber DOI

Abstrak

Research in Natural Language Processing (NLP) and computational linguistics highly depends on a good quality representative corpus of any specific language. Bangla is one of the most spoken languages in the world but Bangla NLP research is in its early stage of development due to the lack of quality public corpus. This article describes the detailed compilation methodology of a comprehensive monolingual Bangla corpus, KUMono (Khulna University Monolingual corpus). The newly developed corpus consists of more than 350 million word tokens and more than one million unique tokens from 18 major text categories of online Bangla websites. We have conducted several word-level and character-level linguistic phenomenon analyses based on empirical studies of the developed corpus. The corpus follows Zipf’s curve and hapax legomena rule. The quality of the corpus is also assessed by analyzing and comparing the inherent sparseness of the corpus with existing Bangla corpora, by analyzing the distribution of function words of the corpus and vocabulary growth rate. We have developed a Bangla article categorization application based on the KUMono corpus and received compelling results by comparing to the state-of-the-art models.

Topik & Kata Kunci

Computer Science

Penulis (7)

A. Akther

Md. Shymon Islam

Hafsa Sultana

A. K. Z. Rasel Rahman

Sujan Kumar Saha

Kazi Masudul Alam

Rameswar Debnath

Format Sitasi

APA MLA BibTeX

Akther, A., Islam, M.S., Sultana, H., Rahman, A.K.Z.R., Saha, S.K., Alam, K.M. et al. (2022). Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono. https://doi.org/10.1109/access.2022.3195236

Akses Cepat

Lihat di Sumber doi.org/10.1109/access.2022.3195236

Informasi Jurnal

Tahun Terbit: 2022
Bahasa: en
Total Sitasi: 16×
Sumber Database: Semantic Scholar
DOI: 10.1109/access.2022.3195236
Akses: Open Access ✓