Published June 7, 2024 | Version v2

Blood Donation Japanese Tweets - Dehydrated versions

Description

These datasets contains tweets collected through Twitter API with the previously available Research Account, during May 2023. Data was collected by using the  “(献血 OR kenketsu OR けんけつ OR ラブラッド OR LoveBlood OR #献血 OR #けんけつ OR #kenketsu OR #ラブラッド OR #LoveBlood) lang:ja” search string, which containts the japanese words related to Blood Donation in japanese, and the name of the official blood donation application from Japan, LoveBlood.

The Twitter_labeled_dataset contains tweets from the year 2022 randomly selected to prepare for manual labeling, while the other two datasets contain the tweets collected from October 2022 to April 2023 for automatized classification. Dataset_unlabeled_2023_full references to the raw collected tweets, while Data_classified_full has the final group of tweets after preprocessing and additional filtering.

Considering the privacy recommendations when using Twitter data, all datasets has been "dehydrated," meaning they only contains the tweet IDs and not the associated content. To use this data, you will need to reassociate the IDs with their corresponding data (rehydration).

Notes

Corrected format for tweet_id column and added the missing sentiment column to the full dataset

Files

data_classified_full.csv

Files (29.9 MB)

Name Size Download all
md5:e623ecf4060e6e528bb8390876d97ecc
19.9 MB Preview Download
md5:75af28be0937699b2707a952f74f31a7
9.8 MB Preview Download
md5:7c5e53c684e7f01a27114308a9f34d8b
184.3 kB Preview Download

Additional details

Related works

Is new version of
Conference proceeding: 10.1109/IEEECONF58974.2023.10404778 (DOI)