Dataset Open Access

Webis Cross-Lingual Sentiment Dataset 2010 (Webis-CLS-10)

Prettenhofer, Peter; Stein, Benno

The Cross-Lingual Sentiment (CLS) dataset comprises about 800.000 Amazon product reviews in the four languages English, German, French, and Japanese.

For more information on the construction of the dataset see (Prettenhofer and Stein, 2010) or the enclosed readme files. If you have a question after reading the paper and the readme files, please contact Peter Prettenhofer.

We provide the dataset in two formats: 1) a processed format which corresponds to the preprocessing (tokenization, etc.) in (Prettenhofer and Stein, 2010); 2) an unprocessed format which contains the full text of the reviews (e.g., for machine translation or feature engineering).

The dataset was first used by (Prettenhofer and Stein, 2010). It consists of Amazon product reviews for three product categories---books, dvds and music---written in four different languages: English, German, French, and Japanese. The German, French, and Japanese reviews were crawled from Amazon in November, 2009. The English reviews were sampled from the Multi-Domain Sentiment Dataset (Blitzer et. al., 2007). For each language-category pair there exist three sets of training documents, test documents, and unlabeled documents. The training and test sets comprise 2.000 documents each, whereas the number of unlabeled documents varies from 9.000 - 170.000.

Files (555.6 MB)
Name Size
cls-acl10-processed-README.txt
md5:57a67c76834dabad00de1514723da8d3
4.5 kB Download
cls-acl10-processed.tar.gz
md5:3d5d7b4fe3945e0f5cf438dd15e1c25d
240.9 MB Download
cls-acl10-unprocessed-README.txt
md5:c37e533660d4e3e42073025bef10a403
4.7 kB Download
cls-acl10-unprocessed.tar.gz
md5:3956bae48add21162ffefb26ae19b266
314.7 MB Download
  • Peter Prettenhofer and Benno Stein. Cross-Language Text Classification using Structural Correspondence Learning. In 48th Annual Meeting of the Association of Computational Linguistics (ACL 10), pages 1118-1127, July 2010. Association for Computational Linguistics

90
49
views
downloads
All versions This version
Views 9090
Downloads 4949
Data volume 7.8 GB7.8 GB
Unique views 4343
Unique downloads 2222

Share

Cite as