Published September 20, 2025
| Version v1
Dataset
Open
A Tale of Two Hallucinations: Empirical Evidence for Distinct Failure Modes in Large Language Models Dataset
Authors/Creators
Description
This dataset is a comprehensive test suite designed to empirically validate and evaluate two distinct types of AI hallucinations: Type 1 (Uncertainty-driven) and Type 2 (Confidence-driven). The data was collected from controlled experiments involving six model families (including Perplexity, Claude Sonnet 4, ChatGPT 5, Deepseek, ChatGPT 4o, and Gemini 2.5 Pro) and a total of 120 test cases.
Files
Files
(102.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:8a6b4f800bf05c97d06ac715369a5d67
|
102.5 kB | Download |
Additional details
Related works
- Is described by
- Preprint: 10.5281/zenodo.17167175 (DOI)