There is a newer version of the record available.

Published April 30, 2025 | Version v1

JailFact-Bench: A Comprehensive Analysis of Jailbreak Attacks vs. Hallucinations in LLMs

  • 1. ROR icon New York University Abu Dhabi
  • 2. New York University Abu Dhabi (NYUAD)
  • 3. ROR icon Ruhr University Bochum

Description

JailFact-Bench is a curated benchmark dataset for analyzing jailbreak attacks and hallucination patterns in Large Language Models (LLMs). It contains semantically aligned jailbreak and factuality prompts, along with metadata including toxicity shifts, similarity scores, and annotation strategies. Developed as part of a capstone research project at NYU Abu Dhabi under Professor Christina Pöpper, this dataset accompanies the paper accepted at the SiMLA 2025 Workshop, co-located with the 23rd International Conference on Applied Cryptography and Network Security (ACNS).

Files

README.md

Files (24.6 kB)

Name Size Download all
md5:ca3ba687e368502983154534ab967a23
22.1 kB Download
md5:d24e7f3390b53637810a8b856201f8d6
2.5 kB Preview Download

Additional details

Dates

Created
2025-04-30
Dataset creation and submission date