Published December 21, 2025 | Version v1

Leveraging Large Language Models for Rare Disease Named Entity Recognition

Authors/Creators

Description

This repository keeps the RareDis Corpus (rds files) and Orphanet (csv file) datasets analyzed in the paper "Leveraging Large Language Models for Rare Disease Named Entity Recognition". The raw datasets are publicly available from the NLP4RARE-CM-UC3M repository at https://github.com/isegura/NLP4RARE-CM-UC3M and Orphadata at https://www.orphadata.com/orphanet-scientific-knowledge/.

Files

orphar.csv

Files (6.5 MB)

Name Size Download all
md5:dee2c1e40bec025341c0c5b11511454c
5.9 MB Preview Download
md5:b94f375d4db3ab93c3364fdf7b9a9efa
53.4 kB Download
md5:bc1364da810255c3227fd945e9bcda67
107.3 kB Download
md5:b681bae76d3286d799a5e56b115964de
368.1 kB Download