Published October 2, 2024 | Version v1

Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code

Authors/Creators

Description

Despite their success, large language models (LLMs) face the critical challenge of hallucinations, generating plausible but incorrect content. While much research has focused on hallucinations in multiple modalities including images and natural language text, less attention has been given to hallucinations in source code, which leads to incorrect and vulnerable code that causes significant financial loss. To pave the way for research in LLMs' hallucinations in code, we introduce Collu-Bench, a benchmark for predicting code hallucinations of LLMs across code generation (CG) and automated program repair (APR) tasks. Collu-Bench includes 13,234 code hallucination instances collected from five datasets and 11 diverse LLMs, ranging from open-source models to commercial ones. 
To better understand and predict code hallucinations, Collu-Bench provides detailed features such as the per-step log probabilities of LLMs' output, token types, and the execution feedback of LLMs' generated code for in-depth analysis. In addition, we conduct experiments to predict hallucination on Collu-Bench, using both traditional machine learning techniques and neural networks, which achieves 22.03 -- 33.15% accuracy. Our experiments draw insightful findings of code hallucination patterns, reveal the challenge of accurately localizing LLMs' hallucinations, and highlight the need for more sophisticated techniques.

Files

dataset_info.json

Files (2.2 GB)

Name Size
md5:7ffdbc2f564f6976a74c6e55e5dd36a9
775.7 MB Download
md5:2d1e873e3c559af0cbcfaeb65461d9b3
833.1 MB Download
md5:a17df4878620a5e12de8c7f73b2a97f1
596.6 MB Download
md5:843462796eaa7ea32b02e119862cf208
3.8 kB Preview Download
md5:69867e2f47083d0751df500d4abcd6ad
98 Bytes Preview Download
md5:5f2695b6c81d315546721c45542bce76
365 Bytes Preview Download