Eighth International Conference on Language Resources and Evaluation | |
A contrastive review of paraphrase acquisition techniques | |
Houda Bouamor ; Aurélien Max ; Gabriel Illouz ; Anne Vilnat | |
Others : http://www.lrec-conf.org/proceedings/lrec2012/pdf/555_Paper.pdf PID : 51611 |
|
来源: CEUR | |
【 摘 要 】
This paper addresses the issue of what approach should be used for building a corpus of sentential paraphrases depending on one’s requirements. Six strategies are studied: (1) multiple translations into a single language from another language; (2) multiple translations into a single language from different other languages; (3) multiple descriptions of short videos; (4) multiple subtitles for the same language; (5) headlines for similar news articles; and (6) sub-sentential paraphrasing in the context of a Web-based game. We report results on French for 50 paraphrase pairs collected for all these strategies, where corpora were manually aligned at the finest possible level to define oracle performance in terms of accessible sub-sentential paraphrases. The differences observed will be used as criteria for motivating the choice of a given approach before attempting to build a new paraphrase corpus.
【 预 览 】
Files | Size | Format | View |
---|---|---|---|
A contrastive review of paraphrase acquisition techniques | 388KB | download |