Comparison of CRF and BiLSTM Models for Disease Named Entity Recognition on the NCBI Disease Dataset
DOI:
https://doi.org/10.18495/comengapp.v15i3.1352Keywords:
Named Entity Recognition, CRF, BiLSTM, NCBI Disease, Biomedical, NLPAbstract
Named Entity Recognition (NER) is a fundamental task in Natural Language Processing (NLP) aimed at identifying and classifying named entities in text. In the biomedical domain, disease entity recognition is crucial for supporting automated medical information extraction. This study compares two NER model approaches, Conditional Random Field (CRF) and Bidirectional Long Short-Term Memory (BiLSTM), for disease entity recognition using the NCBI Disease dataset. The CRF model utilizes handcrafted linguistic features including prefixes, suffixes, POS tags, and word shapes, while BiLSTM employs GloVe 200-dimensional embeddings and automatically learns features from sequential context. BiLSTM experiments were conducted using 5 different seeds for robust statistical validation, reporting mean F1-score, standard deviation, and 95% confidence interval. Comprehensive analysis includes confusion matrix, error analysis, learning curve, and performance comparison per entity length. Results show that CRF achieves an F1-score of 78.49% (P=82.37%, R=74.95%), while BiLSTM achieves 76.54% ± 1.32% (95% CI ±1.15%). CRF outperforms BiLSTM by 1.95 percentage points under baseline conditions, although BiLSTM demonstrates notably fewer false positives (FP=124 vs FP=153) and a standard deviation of 1.32% across 5 experimental seeds
Downloads
Submitted
Accepted
Published
Issue
Section
License
Copyright (c) 2026 Agus Siswanto, Bambang Tutuko, Firdaus, Jasmir

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.







