Published August 12, 2026 | Version 1.0.0

ANERD: A Large-Scale Arabic Corpus for Named Entity Recognition and Disambiguation

  • 1. Department of Computer Science, College of Computer and Information Sciences, Imam Mohammad Ibn Saud Islamic University, Riyadh 11564, Saudi Arabia
  • 2. Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh 11362, Saudi Arabia

Description

ANERD v1.0.0 is a large-scale Arabic corpus designed to support both Named Entity Recognition (NER) and Named Entity Disambiguation/Entity Linking (NED/EL). The corpus contains 502,756 sentences and 11,551,982 tokens. Named entities are annotated using the BIO scheme with four entity types: PERS, ORG, LOC, and MISC. Entity mentions are linked, when applicable, to Arabic Wikipedia through normalized entity titles and URLs; mentions without an appropriate knowledge-base target are marked as NIL. The corpus is organized into Easy, Medium, and Hard sentence-level difficulty categories based on annotation agreement. Each difficulty category is further divided into training, validation, and test subsets using a 60%/20%/20% split. The dataset is distributed as UTF-8 plain-text files in a sentence-delimited, token-per-line format.

Files

ANERD_v1.0.0.zip

Files (48.2 MB)

Name Size Download all
md5:9ebd3e430ec2ae5fb54afa63d16058cb
48.2 MB Preview Download

Additional details

Related works

Is described by
Journal: 10.3390/data11090243 (DOI)

References