Published December 21, 2025
| Version v1
Dataset
Open
Leveraging Large Language Models for Rare Disease Named Entity Recognition
Authors/Creators
Description
This repository keeps the RareDis Corpus (rds files) and Orphanet (csv file) datasets analyzed in the paper "Leveraging Large Language Models for Rare Disease Named Entity Recognition". The raw datasets are publicly available from the NLP4RARE-CM-UC3M repository at https://github.com/isegura/NLP4RARE-CM-UC3M and Orphadata at https://www.orphadata.com/orphanet-scientific-knowledge/.