Published July 21, 2026 | Version v4

Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers

Authors/Creators

  • 1. ROR icon Iran University of Science and Technology

Description

The Noor-Sharaye dataset is a morphologically annotated Classical Arabic corpus containing approximately 205,000 word instances extracted from 17 different Classical Arabic books, including Quranic, Fiqh, Hadith, and historical texts. Each token is enriched with detailed linguistic annotations such as

  • Stem
  •  Lemma
  • Root
  • part-of-speech tags (pos)
  •  Segmentation
  •  Grammatical case
  • Gender, Number
  • Affix-level features

The data are encoded in UTF-8 XLS, XML, and JSON formats for broad compatibility. 

This resource supports stemming, root extraction, morphological analysis, and benchmarking of AI-based models in Arabic Natural Language Processing

Files

metadata.txt

Files (8.3 MB)

Name Size Download all
md5:edc12b8e25610ae02345b9f138943790
1.2 kB Preview Download
md5:7e2e32776efce8569c1ad4f8b3c889e6
421.6 kB Download
md5:86b309c08250c99290fa618999a91f4d
3.5 MB Preview Download
md5:80d2bc1e31c31487a510e3ee331ff40b
4.4 MB Download
md5:2d6176bd0b964d4d742d6d2462f7cbbe
2.4 kB Preview Download

Additional details