This repo explores token classification for abbreviation and long-form detection using RoBERTa. We evaluate the impact of adding 50% of the PLODv2-filtered dataset, achieving improved F1 and recall. The repo includes methodology, evaluation using seqeval, and confusion matrix analysis.
named-entity-recognition sequence-labeling roberta long-form huggingface-transformers token-classification huggingface-datasets biomedical-nlp seqeval abbreviation-detection
-
Updated
Jun 7, 2025 - Python