期刊论文

【摘要】

Bilingual corpora, containing the same documents in two different languages, are becoming an essential resource for natural language processing. Clustering bilingual corpora provides us with an insight into the differences between languages when term frequency-based Information Retrieval (IR) tools are used. It also allows one to use the Natural Language Processing (NLP) and IR tools in one language to implement IR for another language. This study reports on our work on applying Hierarchical Agglomerative Clustering (HAC) to a large corpus of documents where each appears both in Malay and English languages. These documents are clustered for each language and both results are compared with respect to the content of clusters produced. Further, the effects of using different methods of computing the inter-clusters distance on the cluster results is also studied. These methods include Single, Complete and Average links. Finally, this study describes an experiment employing a genetic algorithm to fine-tune individual term’s weight in order to reproduce more closely a predefined set of clusters. In this way, clustering becomes a supervised learning technique that is trained to better reproduce known clusters in Malay language when applied to the corresponding documents in English language. On the data available, the results of clustering one language resemble the other, provided the number of clusters required is relatively small. The method used to compute the inter-clusters distance also influences the cluster results. The result actually showed an increase in the percentage of aligned clusters, when we applied the genetic algorithm to fine-tune weights of terms considered in clustering the bilingual Malay-English corpora. This study concludes that with a smaller number of clusters, k = 5, all of the clusters from English texts can be mapped into the clusters of Malay texts, by using the Complete link distance measure in clustering the bilingual parallel corpus. In contrast, with a large size of clusters, fewer clusters from English texts can be mapped into the clusters of Malay texts.

【授权许可】

Unknown

【预览】

附件列表
Files	Size	Format	View
RO201911300110293ZK.pdf	177KB	PDF	download

Journal of Computer Science
OPTIMIZING CLUSTERS ALIGNMENT FOR BILINGUAL MALAY-ENGLISH CORPORA \| Science Publications

Asni Tahir¹ Chan Chen Jie¹ Ng Zhen Wei¹ Joe Henry Obit¹ Rayner Alfred¹
关键词: Bilingual Corpora; Hierarchical Agglomerative Clustering; Parallel Clustering; Genetic Algorithm; Malay-English Corpora; Knowledge Management;
DOI : 10.3844/jcssp.2012.1970.1978
学科分类：计算机科学（综合）
来源: Science Publications
PDF


	文献评价指标
	下载次数：19次	浏览次数：12次

【 摘 要 】

【 授权许可】

【 预 览 】

【摘要】

【授权许可】

【预览】