Published December 25, 2022
| Version v1
Dataset
Open
Parallel Urdu and Roman Urdu Toxic Comments and Transliteration Corpus
Authors/Creators
Contributors
Annotator:
Description
This dataset is significant for its role in addressing the classification of toxic comments in Urdu. This dataset also contains equivalent Roman Urdu comments for transliteration purpos.
Files
Files
(9.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:8b1ac42391b332d32a51fd06e4e439e4
|
9.2 MB | Download |
Additional details
Additional titles
- Alternative title (Urdu)
- PURUTT
- Alternative title (Urdu)
- PURUTT Corpus
- Alternative title (Urdu)
- Parallel Urdu and Roman Urdu Corpus for Toxic Comments and Transliteration
- Alternative title (Urdu)
- Parallel Urdu and Roman Urdu Toxic Comments and Transliteration
Related works
- Is part of
- Thesis: Detection of Hate Speech in Urdu and Roman Urdu Texts (Other)
- Is published in
- Journal article: 10.1109/ACCESS.2025.3535862 (DOI)
Software
- Repository URL
- https://github.com/hafizhassaan/Urdu-Toxic-Comments.git
- Programming language
- Python
- Development Status
- Active