There is a newer version of the record available.

Published March 2, 2026 | Version 2.0.0

Empirica: Epistemic Self-Assessment for AI Systems — Dataset and Paper v2.0.0

  • 1. Empirica

Description

This dataset contains 1,190,265 evidence observations from 792 AI sessions, capturing epistemic self-assessment across 13 vectors. The data supports the empirical findings in the Empirica paper, demonstrating that effective AI capability genuinely increases during task execution through externalized epistemic scaffolding (the "learning delta").

Key findings from v2.0.0: 83.3% of 456 clean learning pairs showed knowledge improvement with mean capability growth of +0.162. Calibration variance drops 107× as evidence accumulates. Dual-track calibration: 1,190 grounded beliefs with 288 post-test verifications across 9 vectors. Primary AI: Claude (Anthropic). Collection period: August 2025 - March 2026.

Includes: empirica-dataset-v2.0.0.zip (CSV exports: bayesian_beliefs, epistemic_snapshots, cascades, sessions, grounded_beliefs, grounded_verifications, calibration_trajectory, calibration_summary) and empirica-paper-v2.0.0.pdf (54-page paper).

Notes

v2.0.0 update: 13.5× more evidence observations (1,190,265 vs 87,871), grounded verification track (9/13 vectors), dual-track Bayesian calibration, externalized scaffolding framing, 7-layer artifact taxonomy, 4-layer memory architecture. Paper expanded from 36 to 54 pages.

Files

empirica-dataset-v2.0.0.zip

Files (1.7 MB)

Name Size Download all
md5:5ad5e1d7f7469a1ec6efad1914f6f11e
828.1 kB Preview Download
md5:7a08ddb7cca21f416325f48d247acfcf
884.5 kB Preview Download

Additional details

Related works