Published September 26, 2026 | Version prereg-v1

Preregistration: A bias-corrected, entropy-stratified test of whether TypeSafe AI's Jev lowers its confidence when humans disagree (ChaosNLI)

Description

Preregistration for an independent study of TypeSafe AI's Jev decision model (jev-1.13.0). Independent study. Not affiliated with or endorsed by TypeSafe AI.

Question: does Jev's calibration error, scored against the share of 100 ChaosNLI annotators who chose the model's label, increase on items where annotators disagree? Design: ChaosNLI (SNLI and MNLI), 750 items from the lowest quartile of annotator entropy versus 750 from the highest (equal n). Primary endpoint: bias-corrected difference in top-label expected calibration error (hard minus easy), with a parametric Monte Carlo test. Verdict rule: "tracks" if corrected ΔECE ≥ 0.09 and p < 0.05; "holds" if ΔECE ≥ 0.09 is rejected one-sided and corrected ΔECE ≤ 0.02; otherwise "inconclusive". Results were committed to be published regardless of direction.

Timing. This record is deposited after data collection. The evidence that the plan preceded the data is independent of this deposit: the frozen plan and its checksums were submitted to OpenTimestamps at 2026-09-25 06:29:11 UTC (proofs included as .ots files, anchored in Bitcoin, verifiable with "ots verify"), and the scored run began at 07:01:59 UTC the same day.

Identifiers
Frozen commit: 3b7bcabf6c4ef71adc95e7e5546d3073612c2e88 (tag prereg-v1-frozen)
Binding commit, Amendment 11: ed4655f9cc5dbf0bd7ea90b5929a72794d5a9495 (tag prereg-v1-binding)
SHA-256 of PREREGISTRATION.md: e19aecb696d780d16d05363b34d850b9a58282097cd84b0b207df7eb77146014
Repository: https://github.com/GautamTalksDev/jevbench

Files: PREREGISTRATION.md (the plan and all amendments); PREREG_CHANGELOG.md (every edit after the binding commit, classified; all clerical); preregistration.lock.json (dataset hashes); jevbench-prereg-v1-frozen.zip (repository snapshot at the frozen commit); SHA256SUMS.txt; two OpenTimestamps proofs.

Notes. ChaosNLI sentence text is not included (CC BY-NC 4.0); it is fetched from the original source by scripts/fetch_chaosnli.py. runs/offline_fixture in the snapshot is a synthetic demonstration fixture added on 2026-09-23, before any access to Jev; its model field is a schema placeholder and it contains no real API responses. Code was written with AI assistance (Cursor and Claude); the author reviewed the design, code changes and all hashes.

Files

jevbench-prereg-v1-frozen.zip

Files (2.7 MB)

Name Size Download all
md5:c4791e72a3fc0c20591173785d431439
2.7 MB Preview Download
md5:145789accb8b20ba35f0af0381be2b83
4.1 kB Preview Download
md5:8a4d437c3173759a9833696a5da69178
563 Bytes Preview Download
md5:0a576f189ec1364b6241fa01a2ecbf75
29.2 kB Preview Download
md5:695712b6c95c3b723201916866f22902
3.9 kB Download
md5:053ab7d59151b5f6f90aa32e39b36d98
359 Bytes Preview Download
md5:d408ec0fe0455b3f383d14db636dd746
3.9 kB Download

Additional details

Related works

Is referenced by
Preprint: 10.5281/zenodo.22971492 (DOI)
Is supplemented by
Software: https://github.com/GautamTalksDev/jevbench (URL)
References
Conference paper: 10.18653/v1/2020.emnlp-main.734 (DOI)

Software

Repository URL
https://github.com/GautamTalksDev/jevbench
Programming language
Python
Development Status
Active