DOAJ Open Access 2020

PAAD: POLITICAL ARABIC ARTICLES DATASET FOR AUTOMATIC TEXT CATEGORIZATION

Dhafar Hamed Abd Ahmed T. Sadiq Ayad R. Abbas

Abstrak

Now day’s text Classification and Sentiment analysis is considered as one of the popular Natural Language Processing (NLP) tasks. This kind of technique plays significant role in human activities and has impact on the daily behaviours. Each article in different fields such as politics and business represent different opinions according to the writer tendency. A huge amount of data will be acquired through that differentiation. The capability to manage the political orientation of an online article automatically. Therefore, there is no corpus for political categorization was directed towards this task in Arabic, due to the lack of rich representative resources for training an Arabic text classifier. However, we introduce political Arabic articles dataset (PAAD) of textual data collected from newspapers, social network, general forum and ideology website. The dataset is 206 articles distributed into three categories as (Reform, Conservative and Revolutionary) that we offer to the research community on Arabic computational linguistics. We anticipate that this dataset would make a great aid for a variety of NLP tasks on Modern Standard Arabic, political text classification purposes. We present the data in raw form and excel file. Excel file will be in four types such as V1 raw data, V2 preprocessing, V3 root stemming and V4 light stemming.

Topik & Kata Kunci

Penulis (3)

D

Dhafar Hamed Abd

A

Ahmed T. Sadiq

A

Ayad R. Abbas

Format Sitasi

Abd, D.H., Sadiq, A.T., Abbas, A.R. (2020). PAAD: POLITICAL ARABIC ARTICLES DATASET FOR AUTOMATIC TEXT CATEGORIZATION. https://doi.org/10.25195/ijci.v46i1.246

Akses Cepat

PDF tidak tersedia langsung

Cek di sumber asli →
Lihat di Sumber doi.org/10.25195/ijci.v46i1.246
Informasi Jurnal
Tahun Terbit
2020
Sumber Database
DOAJ
DOI
10.25195/ijci.v46i1.246
Akses
Open Access ✓