ANERD: A Large-Scale Arabic Corpus for Named Entity Recognition and Disambiguation
- 1. Department of Computer Science, College of Computer and Information Sciences, Imam Mohammad Ibn Saud Islamic University, Riyadh 11564, Saudi Arabia
- 2. Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh 11362, Saudi Arabia
Description
ANERD v1.0.0 is a large-scale Arabic corpus designed to support both Named Entity Recognition (NER) and Named Entity Disambiguation/Entity Linking (NED/EL). The corpus contains 502,756 sentences and 11,551,982 tokens. Named entities are annotated using the BIO scheme with four entity types: PERS, ORG, LOC, and MISC. Entity mentions are linked, when applicable, to Arabic Wikipedia through normalized entity titles and URLs; mentions without an appropriate knowledge-base target are marked as NIL. The corpus is organized into Easy, Medium, and Hard sentence-level difficulty categories based on annotation agreement. Each difficulty category is further divided into training, validation, and test subsets using a 60%/20%/20% split. The dataset is distributed as UTF-8 plain-text files in a sentence-delimited, token-per-line format.
Files
ANERD_v1.0.0.zip
Files
(48.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:9ebd3e430ec2ae5fb54afa63d16058cb
|
48.2 MB | Preview Download |
Additional details
Related works
- Is described by
- Journal: 10.3390/data11090243 (DOI)
References
- Alotaibi, M. S.; Menai, M. E. B. ANERD: A Large-Scale Arabic Corpus for Named Entity Recognition and Disambiguation. Data 2026, 11, 243. https://doi.org/10.3390/data11090243