Published August 4, 2026 | Version v1.4

AIFaultBench: A Reproducible Benchmark of Real-World AI Software Faults

  • 1. Dalhousie University
  • 2. Polytechnique MontrĂ©al

Description

AIFaultBench is a benchmark of 770 real-world AI software faults collected from 105 open-source repositories across 76 organizations, spanning traditional machine learning, deep learning, large language model infrastructure, reinforcement learning, agentic AI systems, and AI tooling. Each fault ships with the original GitHub issue report, a minimal reproduction script, a dependency specification, codebase reconstruction and environment setup scripts, reproduction logs, structured metadata, and a reproduction trajectory. 652 of the 770 faults (85%) are verified reproducible; the remainder document the reasons preventing reproduction.

Notes

If you use AIFaultBench, please cite it as below.

Files

mehilshah/AIFaultBench-v1.4.zip

Files (18.5 MB)

Name Size Download all
md5:894a9211d03a20bf1bdac502c3bf0258
18.5 MB Preview Download

Additional details

Related works