Published May 15, 2026 | Version 1.1.0

UGSC Multilingual Sentiment Dataset for Sustainable Urban Mobility

Authors/Creators

Description

The UGSC-ML dataset is the multilingual version of the User Gold Standard Corpus (UGSC)
for sustainable urban mobility sentiment analysis.

It consists of 375 English transport-related user reviews from TripAdvisor with
sentence-aligned translations in Spanish, French, German, and Italian, yielding
1,875 instances across five languages.

The CardiffNLP XLM-RoBERTa model (`cardiffnlp/twitter-xlm-roberta-base-sentiment`)
is applied in zero-shot domain transfer mode: pre-trained on CC100 and fine-tuned
on Twitter sentiment data, then applied without any task- or domain-specific fine-tuning
to transport reviews.

Files

low_confidence_annotated_113.csv

Files (771.6 kB)

Name Size Download all
md5:f26d753ac2c95da3383400531ae3d2bc
33.9 kB Preview Download
md5:b027a7cfb637de972c8f1acbfc35d8cc
5.7 kB Preview Download
md5:13a78d74422dd5d5584fb8e7eee94148
394.4 kB Preview Download
md5:f2f8be278be6c63db37b79c34aa7c0bf
337.6 kB Preview Download