期刊论文

【摘要】

The development of natural language processing resources for Albanian has grown steadily in recent years. This paper presents research conducted on unsupervised learning-the challenges associated with building a dictionary for the Albanian language and creating part-of-speech tagging models. The majority of languages have their own dictionary, but languages with low resources suffer from a lack of resources. It facilitates the sharing of information and services for users and whole communities through natural language processing. The experimentation corpora for the Albanian language includes 250K sentences from different disciplines, with a proposal for a part-of-speech tagging tag set that can adequately represent the underlying linguistic phenomena. Contributing to the development of Albanian is the purpose of this paper. The results of experiments with the Albanian language corpus revealed that its use of articles and pronouns resembles that of more high-resource languages. According to this study, the total expected frequency as a means for correctly tagging words has been proven effective for populating the Albanian language dictionary.

【授权许可】

CC BY

【预览】

附件列表
Files	Size	Format	View
RO202306300002670ZK.pdf	455KB	PDF	download

Annals of Emerging Technologies in Computing
Building Dictionaries for Low Resource Languages: Challenges of Unsupervised Learning
article
Mati, Diellza Nagavci¹ Hamiti, Mentor¹ Susuri, Arsim² Selimi, Besnik¹ Ajdari, Jaumin¹
[1] South East European University;University of Prizren
关键词: Albanian language; corpora; dictionaries; natural language processing; part-of-speech tagging;
DOI : 10.33166/AETiC.2021.03.005
学科分类：电子与电气工程
来源: International Association for Educators and Researchers (IAER)
PDF


	文献评价指标
	下载次数：9次	浏览次数：3次

【 摘 要 】

【 授权许可】

【 预 览 】

【摘要】

【授权许可】

【预览】