Human vs ChatGPT: Effect of Data Annotation in Interpretable Crisis-Related Microblog Classification

Nguyen, Thi Huyen

doi:10.1145/3589334.3648141

Published May 13, 2024 | Version v1

Conference paper Open

Human vs ChatGPT: Effect of Data Annotation in Interpretable Crisis-Related Microblog Classification

Nguyen, Thi Huyen¹

1. L3S Research Center

Recent studies have exploited the vital role of microblogging platforms, such as Twitter, in crisis situations. Various machine-learning approaches have been proposed to identify and prioritize crucial information from different humanitarian categories for preparation and rescue purposes. In crisis domain, the explanation of models' output decisions is gaining significant research momentum. Some previous works focused on human annotations of rationales to train and extract supporting evidence for model interpretability. However, such annotations are usually expensive, require much effort, and are not always available in real-time situations of a new crisis event. In this paper, we investigate the recent advances in large language models (LLMs) as data annotators on informal tweet text. We perform a detailed qualitative and quantitative evaluation of ChatGPT rationale annotations over a few-shot setup. ChatGPT annotations are quite close to humans but less precise in nature. Further, we propose an active learning-based interpretable classification model from a small set of annotated data. Our experiments show that (a). ChatGPT has the potential to extract rationales for the crisis tweet classification tasks, but the performance is slightly less than the model trained on human-annotated rationale data (\sim3-6%), (b). active learning setup can help reduce the burden of manual annotations and maintain a trade-off between performance and data size.

Files

3589334.3648141.pdf

Files (1.4 MB)

Name	Size	Download all
3589334.3648141.pdf md5:98d0e687c22b0410f1186de7ac67e332	1.4 MB	Preview Download

Additional details

Accepted: 2024-05-13

	All versions	This version
Views	71	71
Downloads	47	47
Data volume	69.2 MB	69.2 MB

Human vs ChatGPT: Effect of Data Annotation in Interpretable Crisis-Related Microblog Classification

Creators

Description

Files

3589334.3648141.pdf

Files (1.4 MB)

Additional details

Dates