期刊论文详细信息
BMC Bioinformatics
A new synonym-substitution method to enrich the human phenotype ontology
Research Article
Ranga C. Gudivada1  Diego Martinez2  Maria Taboada3  Hadriana Rodriguez3 
[1] CareCentrix, 06103, Hartford, Conneticut, USA;Department of Applied Physics, University of Santiago de Compostela, 15705, Campus Vida, Santiago de Compostela, Spain;Department of Electronics & Computer Science, University of Santiago de Compostela, Campus Vida, 15705, Santiago de Compostela, Spain;
关键词: Biomedical ontologies;    Entity name discovery;    Human phenotype ontology;    PubMed;   
DOI  :  10.1186/s12859-017-1858-7
 received in 2017-06-23, accepted in 2017-10-02,  发布年份 2017
来源: Springer
PDF
【 摘 要 】

BackgroundNamed entity recognition is critical for biomedical text mining, where it is not unusual to find entities labeled by a wide range of different terms. Nowadays, ontologies are one of the crucial enabling technologies in bioinformatics, providing resources for improved natural language processing tasks. However, biomedical ontology-based named entity recognition continues to be a major research problem.ResultsThis paper presents an automated synonym-substitution method to enrich the Human Phenotype Ontology (HPO) with new synonyms. The approach is mainly based on both the lexical properties of the terms and the hierarchical structure of the ontology. By scanning the lexical difference between a term and its descendant terms, the method can learn new names and modifiers in order to generate synonyms for the descendant terms. By searching for the exact phrases in MEDLINE, the method can automatically rule out illogical candidate synonyms. In total, 745 new terms were identified. These terms were indirectly evaluated through the concept annotations on a gold standard corpus and also by document retrieval on a collection of abstracts on hereditary diseases. A moderate improvement in the F-measure performance on the gold standard corpus was observed. Additionally, 6% more abstracts on hereditary diseases were retrieved, and this percentage was 33% higher if only the highly informative concepts were considered.ConclusionsA synonym-substitution procedure that leverages the HPO hierarchical structure works well for a reliable and automatic extension of the terminology. The results show that the generated synonyms have a positive impact on concept recognition, mainly those synonyms corresponding to highly informative HPO terms.

【 授权许可】

CC BY   
© The Author(s). 2017

【 预 览 】
附件列表
Files Size Format View
RO202311098618907ZK.pdf 1897KB PDF download
【 参考文献 】
  • [1]
  • [2]
  • [3]
  • [4]
  • [5]
  • [6]
  • [7]
  • [8]
  • [9]
  • [10]
  • [11]
  • [12]
  • [13]
  • [14]
  • [15]
  • [16]
  • [17]
  • [18]
  • [19]
  • [20]
  • [21]
  • [22]
  • [23]
  • [24]
  • [25]
  • [26]
  • [27]
  • [28]
  • [29]
  • [30]
  • [31]
  • [32]
  • [33]
  • [34]
  • [35]
  • [36]
  文献评价指标  
  下载次数:9次 浏览次数:1次