Published May 15, 2026 | Version v3

A Cannabis Use Reddit Dataset for Aspect-Based Sentiment Analysis

Description

Abstract

Using publicly accessible Reddit posts, we developed a manually annotated dataset for traditional and aspect-based sentiment analysis (ABSA) of cannabis-related discussions in the context of pain management. The dataset consists of 479 post-aspect pairs extracted from specific Reddit communities associated with autoimmune rheumatic diseases (ARDs). We filtered posts using a structured list of cannabis-related terms and extracted context using rule-based sentence segmentation. Subsequently, we manually annotated each post for both traditional and aspect-specific sentiment (positive, negative, neutral). Inter-annotator reliability was assessed using Krippendorff’s ⍺, yielding ⍺ = 0.604 for traditional sentiment and ⍺ = 0.526 for aspect-based sentiment. The dataset offers valuable resources for training, benchmarking, and evaluating machine learning models for ABSA in health-related social media contexts. This dataset can support research in natural language processing, public health informatics, pain medicine, digital epidemiology, and social media-based health monitoring.

 

File structure

  • cannabis_reddit_absa.csv: Dataset containing Reddit post IDs, matched aspect terms, final traditional and aspect-based sentiment labels, and individual annotator labels for each post–aspect pair.
  • reddit_data_pipeline.ipynb: Jupyter Notebook containing functions for sentence segmentation with spaCy, and matching posts to the cannabis term list to generate post–aspect pairs for annotation.
  • cannabis_lexicon.csv: Complete list of cannabis-related terms used to filter and match relevant content in Reddit posts, including abbreviations, synonyms, and common spellings.
  • annotation_guidelines.pdf: Detailed instructions for annotators, including definitions of sentiment labels (neutral, positive, negative, blank), rules for assigning traditional vs. aspect-based labels, guidance for handling personal experience vs. general commentary, and illustrative examples.
  • top_subreddits_weekly_extraction.py: Python script implementing a data extraction and processing workflow for Reddit using PRAW, with outputs stored in MongoDB. Compatibility with newer versions of PRAW is not guaranteed, and the script may require updates to run.

Files

annotation_guidelines.pdf

Files (264.4 kB)

Name Size Download all
md5:a044b1538dd644b0a5c355a0d457ad68
55.1 kB Preview Download
md5:a8b2b3b644b96b1d2eac52e699794a19
161.1 kB Preview Download
md5:51f42a11e3448697c2c70fde779f3761
33.4 kB Preview Download
md5:d7192daafaa1c499a199d53e88f5260c
622 Bytes Preview Download
md5:ced06ba23c507f67113bbdd3b16dc33c
4.5 kB Preview Download
md5:5b666d7953ccd890656fae38b86cdb5e
9.7 kB Download