PHINC: A Parallel Hinglish Social Media Code-Mixed Corpus for Machine Translation

Vivek Srivastava; Mayank Singh

doi:10.5281/zenodo.3605597

Published January 12, 2020 | Version v1

Preprint Open

PHINC: A Parallel Hinglish Social Media Code-Mixed Corpus for Machine Translation

1. Indian Institute of Technology Gandhinagar

Code-mixing is the phenomenon of using more than one language in a sentence. It is a very frequently observed pattern of communication on social media platforms. Flexibility to use mixed languages in one text message might help to communicate efficiently with the target audience. But, it adds to the challenge of processing and understanding natural language to a much larger extent. Here, we are presenting a parallel corpus of the 13,738 code-mixed English-Hindi sentences and their corresponding translation in English. The translations of sentences are done manually by the annotators. We are releasing the parallel corpus to facilitate future research opportunities for code-mixed machine translation.

If you are using this dataset as part of your research, please cite the following paper

@article{srivastava2020phinc,
  title={PHINC: A Parallel Hinglish Social Media Code-Mixed Corpus for Machine Translation},
  author={Srivastava, Vivek and Singh, Mayank},
  journal={arXiv preprint arXiv:2004.09447},
  year={2020}
}

Files

English-Hindi code-mixed parallel corpus.csv

Files (2.1 MB)

Name	Size	Download all
English-Hindi code-mixed parallel corpus.csv md5:8d423e1a34abf6a936c79fe7aa0727d6	2.1 MB	Preview Download

Additional details

URL: https://arxiv.org/abs/2004.09447

	All versions	This version
Views	4,129	4,091
Downloads	262,217	262,204
Data volume	604.9 GB	604.8 GB

PHINC: A Parallel Hinglish Social Media Code-Mixed Corpus for Machine Translation

Authors/Creators

Description

Files

English-Hindi code-mixed parallel corpus.csv

Files (2.1 MB)

Additional details

Identifiers