HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings

doi:10.48550/arXiv.2411.10724

HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings

One of the key tasks in modern applied computational linguistics is constructing word vector representations (word embeddings), which are widely used to address natural language processing tasks such as sentiment analysis, information extraction, and more. To choose an appropriate method for generating these word embeddings, quality assessment techniques are often necessary. A standard approach involves calculating distances between vectors for words with expert-assessed 'similarity'. This work introduces the first 'silver standard' dataset for such tasks in the Kyrgyz language, alongside training corresponding models and validating the dataset's suitability through quality evaluation metrics.

Publication:

arXiv e-prints

Pub Date:

November 2024

DOI:

10.48550/arXiv.2411.10724

arXiv:

arXiv:2411.10724

Bibcode:

2024arXiv241110724A

Keywords:

Computer Science - Computation and Language

E-Print:

Herald of KSTU 68(4) (2023)

NASA/ADS

HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings

Abstract