会议论文详细信息
Eighth International Conference on Language Resources and Evaluation
Annotated Corpora for Word Alignment Between Japanese and English and its Evaluation with MAP-based Word Aligner
Tsuyoshi Okita
Others  :  http://www.lrec-conf.org/proceedings/lrec2012/pdf/1109_Paper.pdf
PID  :  51676
来源: CEUR
PDF
【 摘 要 】

This paper presents two annotated corpora for word alignment between Japanese and English. We annotated on top of the IWSLT-2006 and the NTCIR-8 corpora. The IWSLT-2006 corpus is in the domain of travel conversation while the NTCIR-8 corpus is in the domain of patent. We annotated the first 500 sentence pairs from the IWSLT-2006 corpus and the first 100 sentence pairs from the NTCIR-8 corpus. After mentioned the annotation guideline, we present two evaluation algorithms how to use such hand-annotated corpora: although one is a well-known algorithm for word alignment researchers, one is novel which intends to evaluate a MAP-based word aligner of Okita et al. (2010b).

【 预 览 】
附件列表
Files Size Format View
Annotated Corpora for Word Alignment Between Japanese and English and its Evaluation with MAP-based Word Aligner 625KB PDF download
  文献评价指标  
  下载次数:4次 浏览次数:4次