期刊论文详细信息
BMC Medical Informatics and Decision Making
Qcorp: an annotated classification corpus of Chinese health questions
Haihong Guo1  Xu Na1  Jiao Li1 
[1] Institute of Medical Information / Medical Library, Chinese Academy of Medical Sciences & Peking Union Medical College;
关键词: Health Question;    Annotation;    Classification;    Question Answering;    Chinese;   
DOI  :  10.1186/s12911-018-0593-y
来源: DOAJ
【 摘 要 】

Abstract Background Health question-answering (QA) systems have become a typical application scenario of Artificial Intelligent (AI). An annotated question corpus is prerequisite for training machines to understand health information needs of users. Thus, we aimed to develop an annotated classification corpus of Chinese health questions (Qcorp) and make it openly accessible. Methods We developed a two-layered classification schema and corresponding annotation rules on basis of our previous work. Using the schema, we annotated 5000 questions that were randomly selected from 5 Chinese health websites within 6 broad sections. 8 annotators participated in the annotation task, and the inter-annotator agreement was evaluated to ensure the corpus quality. Furthermore, the distribution and relationship of the annotated tags were measured by descriptive statistics and social network map. Results The questions were annotated using 7101 tags that covers 29 topic categories in the two-layered schema. In our released corpus, the distribution of questions on the top-layered categories was treatment of 64.22%, diagnosis of 37.14%, epidemiology of 14.96%, healthy lifestyle of 10.38%, and health provider choice of 4.54% respectively. Both the annotated health questions and annotation schema were openly accessible on the Qcorp website. Users can download the annotated Chinese questions in CSV, XML, and HTML format. Conclusions We developed a Chinese health question corpus including 5000 manually annotated questions. It is openly accessible and would contribute to the intelligent health QA system development.

【 授权许可】

Unknown   

  文献评价指标  
  下载次数:0次 浏览次数:1次