Semantic Scholar Open Access 2017 10 sitasi

Evaluation of Croatian Word Embeddings

Lukás Svoboda Slobodan Beliga

Lihat Sumber

Abstrak

Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy corpus based on the original English Word2vec word analogy corpus and added some of the specific linguistic aspects from Croatian language. Next, we created Croatian WordSim353 and RG65 corpora for a basic evaluation of word similarities. We compared created corpora on two popular word representation models, based on Word2Vec tool and fastText tool. Models has been trained on 1.37B tokens training data corpus and tested on a new robust Croatian word analogy corpus. Results show that models are able to create meaningful word representation. This research has shown that free word order and the higher morphological complexity of Croatian language influences the quality of resulting word embeddings.

Topik & Kata Kunci

Computer Science

Penulis (2)

Lukás Svoboda

Slobodan Beliga

Format Sitasi

APA MLA BibTeX

Svoboda, L., Beliga, S. (2017). Evaluation of Croatian Word Embeddings. https://www.semanticscholar.org/paper/256ec8b92e0ce41c5d59ceba3ee301deb6621e7a

Akses Cepat

PDF tidak tersedia langsung

Cek di sumber asli →

Lihat di Sumber

Informasi Jurnal

Tahun Terbit: 2017
Bahasa: en
Total Sitasi: 10×
Sumber Database: Semantic Scholar
Akses: Open Access ✓