期刊论文详细信息
BMC Bioinformatics
Parallel sequence tagging for concept recognition
Fabio Rinaldi1  Joseph Cornelius1  Lenz Furrer2 
[1] Dalle Molle Institute for Artificial Intelligence Research (IDSIA USI/SUPSI);Department of Computational Linguistics, University of Zurich;
关键词: Text mining;    Named entity recognition and normalization;    Concept recognition;    Neural network;    Sequence tagging;   
DOI  :  10.1186/s12859-021-04511-y
来源: DOAJ
【 摘 要 】

Abstract Background Named Entity Recognition (NER) and Normalisation (NEN) are core components of any text-mining system for biomedical texts. In a traditional concept-recognition pipeline, these tasks are combined in a serial way, which is inherently prone to error propagation from NER to NEN. We propose a parallel architecture, where both NER and NEN are modeled as a sequence-labeling task, operating directly on the source text. We examine different harmonisation strategies for merging the predictions of the two classifiers into a single output sequence. Results We test our approach on the recent Version 4 of the CRAFT corpus. In all 20 annotation sets of the concept-annotation task, our system outperforms the pipeline system reported as a baseline in the CRAFT shared task, a competition of the BioNLP Open Shared Tasks 2019. We further refine the systems from the shared task by optimising the harmonisation strategy separately for each annotation set. Conclusions Our analysis shows that the strengths of the two classifiers can be combined in a fruitful way. However, prediction harmonisation requires individual calibration on a development set for each annotation set. This allows achieving a good trade-off between established knowledge (training set) and novel information (unseen concepts).

【 授权许可】

Unknown   

  文献评价指标  
  下载次数:0次 浏览次数:5次