Published July 21, 2026
| Version v4
Dataset
Open
Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers
Description
The Noor-Sharaye dataset is a morphologically annotated Classical Arabic corpus containing approximately 205,000 word instances extracted from 17 different Classical Arabic books, including Quranic, Fiqh, Hadith, and historical texts. Each token is enriched with detailed linguistic annotations such as
- Stem
- Lemma
- Root
- part-of-speech tags (pos)
- Segmentation
- Grammatical case
- Gender, Number
- Affix-level features
The data are encoded in UTF-8 XLS, XML, and JSON formats for broad compatibility.
This resource supports stemming, root extraction, morphological analysis, and benchmarking of AI-based models in Arabic Natural Language Processing
Files
metadata.txt
Files
(8.3 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:edc12b8e25610ae02345b9f138943790
|
1.2 kB | Preview Download |
|
md5:7e2e32776efce8569c1ad4f8b3c889e6
|
421.6 kB | Download |
|
md5:86b309c08250c99290fa618999a91f4d
|
3.5 MB | Preview Download |
|
md5:80d2bc1e31c31487a510e3ee331ff40b
|
4.4 MB | Download |
|
md5:2d6176bd0b964d4d742d6d2462f7cbbe
|
2.4 kB | Preview Download |