Published September 4, 2025 | Version v1
Dataset Restricted

Parallel Communities Across the Surface Web and the Dark Web

  • 1. ROR icon Max Planck Institute for Security and Privacy
  • 2. ROR icon Korea Advanced Institute of Science and Technology
  • 3. ROR icon Indian Institute of Technology Delhi

Description

DATA DESCRIPTION

This dataset supports a comparative analysis of online communities across two distinct platforms: Reddit (a mainstream forum) and Dread (a dark web discussion forum). It includes:

  • 7M+ posts and comments from Reddit

  • 200K+ posts and comments from Dread

The data is pre-processed and aligned across similar topics. We provide hashed user identifiers, toxicity scores (via Detoxify), timestamps (in UTC), and cleaned text content. Personally identifiable information has been removed or anonymized.

Field Name Field Description
id A unique identifier for each entry in the dataset.
parent_id Identifier used when the entry is part of a conversation thread or linked to a comment; used to associate replies with their parent.
processed_text The text content of the comment or post, pre-processed (e.g., cleaned, normalized) for analysis.
score / vote The numeric score or vote count associated with the entry.
timestamp The UTC timestamp indicating when the entry was created.
subreddit_name / community_name The name of the community (e.g., subreddit) where the entry was posted.
hashed_user_id A pseudonymized identifier for the user who created the entry, generated using a salted hash. No original usernames are retained.
toxic A numerical score indicating the level of toxicity (https://huggingface.co/unitary/toxic-bert)

 

DATA ACCESS INTRUCTIONS

This dataset is under restricted access. To request access, please follow the steps below:

  1. Login to Zenodo account.

  2. On the Files section, enter your details (email and name).

  3. In the message box state your institutional affiliation and provide a brief description of your intended use of the data.

  4. Click on "Request accces". 

  5. Your request will be reviewed within 3–5 business days. You will receive an email notification once access has been granted.

If you have questions or need support with your request, please contact us at:
megha[dot]sundriyal[at]mpi-sp[dot]org

 

ETHICAL USAGE COMMITMENT

By requesting access, users acknowledge and agree to use the dataset solely for ethical and lawful research purposes. The following uses are strictly prohibited:

  • Developing tools or algorithms intended to promote or assist in illicit activities or to evade law enforcement or regulatory oversight.

  • Any form of algorithmic discrimination or bias based on protected characteristics such as race, gender, sexual orientation, religion, or political beliefs.

  • Unauthorized access, data leakage, or any action that compromises the confidentiality or integrity of the dataset.

Notes

CITATION

To comply with legal and ethical standards, please ensure our work is cited as:

@inproceedings{wenchao2025,
  title={Parallel Communities Across the Surface Web and the Dark Web},
  author={Dong, Wenchao and Sundriyal, Megha and Park, Seongchan and Kim, Jaehong and Cha, Meeyoung and Chakraborty, Tanmoy and Lee, Wonjae},
  booktitle={Findings of the Association for Computational Linguistics: EMNLP 2025},
  year={2025}
}

Files

Restricted

The record is publicly accessible, but files are restricted. Log in to check if you have access.

Request access

If you would like to request access to these files, please fill out the form below.

You are currently not logged in. Do you have an account? Log in here