Published December 29, 2025 | Version 0.1.0

Data-Centric Roman Urdu NLP: High-Quality Dataset Curation, Privacy-Preserving Embeddings, and State-of-the-Art Model Benchmarking

  • 1. Emerson University Multan

Contributors

  • 1. Emerson University Multan

Description

This work presents the creation, benchmarking, and analysis of the largest curated Roman Urdu sentiment dataset to date, accompanied by a fine-tuned, state-of-the-art XLM-R sentiment classifier. It addresses the unique challenges of low-resource languages, including noisy, informal, and slang text, and provides a fully reproducible research artifact for the broader NLP and machine learning community. The dataset and model enable rigorous evaluation, benchmarking, and development of sentiment analysis applications in Roman Urdu, a language variety that has been historically underrepresented in computational research.

Key Contributions:

  • Dataset: 99,000 labeled Roman Urdu text samples collected from WhatsApp chats, Twitter, YouTube comments, and other social media platforms. Data cleaning, deduplication, noise reduction, and class balancing were applied to ensure high quality.

  • Embeddings: Sentence-level embeddings created using XLM-R are provided in 10 chunks of 10,000 rows each, with corresponding labels, supporting reproducible experiments and benchmarking.

  • Model: Fine-tuned XLM-R sentiment classifier achieving state-of-the-art performance on slang and formal text subsets (Validation Accuracy: 0.8157, Macro F1: 0.8148).

  • Benchmarking: Comparative evaluation against existing models demonstrates superior performance, with radar plots and leaderboard bar charts provided in the paper.

  • Discussion: The paper includes an analysis of why dedicated sentiment models outperform general-purpose large language models (LLMs) for this low-resource language task.

  • Reproducibility: All resources are publicly available with persistent DOIs:

This preprint aims to provide a complete and reproducible research artifact, enabling researchers, practitioners, and developers to explore sentiment analysis in low-resource languages, benchmark models, and build downstream applications.

Files

data_centric_roman_urdu_nlp.pdf

Files (238.8 kB)

Name Size Download all
md5:888d683a6e19ba1422d6dc6e098d3ce7
238.8 kB Preview Download

Additional details

Software