学位论文详细信息
Exploitation of Redundant Inverse Term Frequency for Answer Extraction
Computer Science;Question Answering;Information Retrieval;Natural Language Processing;Retrieval Models and Ranking;Query Intent
Lynam, Thomas
University of Waterloo
关键词: Computer Science;    Question Answering;    Information Retrieval;    Natural Language Processing;    Retrieval Models and Ranking;    Query Intent;   
Others  :  https://uwspace.uwaterloo.ca/bitstream/10012/1190/1/trlynam2002.pdf
瑞士|英语
来源: UWSPACE Waterloo Institutional Repository
PDF
【 摘 要 】

An automatic question answering system must find, within a corpus,short factual answers to questions posed in natural language. The process involves analyzing the question, retrieving information related to the question, and extracting answers from the retrieved information. This thesis presents a novel approach to answer extraction in an automated question answering (QA) system. The answer extraction approach is an extension of the MultiText QA system. This system employs a question analysis component to examine the question and to produce query terms for the retrieval component which extracts several document fragments from the corpus. The answer extraction component selects a few short answers from these fragments. This thesis describes the design and evaluation of the Redundant Inverse Term Frequency (RITF) answer extraction component. The RITF algorithm locates and evaluates words from the passages that are likely to be associated with the answer. Answers are selected by finding short fragments of text that contain the most likely words based on: the frequency of the words in the corpus, the number of fragments in which the word occurs, the rank of the passages as determined by the IR, the distance of the word from the centre of the fragment, and category information found through question analysis. RITF makes a substantial contribution in overall results, nearly doubling the Mean Reciprocal Rank (MRR), a standard measure for evaluating QA systems.

【 预 览 】
附件列表
Files Size Format View
Exploitation of Redundant Inverse Term Frequency for Answer Extraction 400KB PDF download
  文献评价指标  
  下载次数:9次 浏览次数:18次