Preregistration: A bias-corrected, entropy-stratified test of whether TypeSafe AI's Jev lowers its confidence when humans disagree (ChaosNLI)
Authors/Creators
Description
Preregistration for an independent study of TypeSafe AI's Jev decision model (jev-1.13.0). Independent study. Not affiliated with or endorsed by TypeSafe AI.
Question: does Jev's calibration error, scored against the share of 100 ChaosNLI annotators who chose the model's label, increase on items where annotators disagree? Design: ChaosNLI (SNLI and MNLI), 750 items from the lowest quartile of annotator entropy versus 750 from the highest (equal n). Primary endpoint: bias-corrected difference in top-label expected calibration error (hard minus easy), with a parametric Monte Carlo test. Verdict rule: "tracks" if corrected ΔECE ≥ 0.09 and p < 0.05; "holds" if ΔECE ≥ 0.09 is rejected one-sided and corrected ΔECE ≤ 0.02; otherwise "inconclusive". Results were committed to be published regardless of direction.
Timing. This record is deposited after data collection. The evidence that the plan preceded the data is independent of this deposit: the frozen plan and its checksums were submitted to OpenTimestamps at 2026-09-25 06:29:11 UTC (proofs included as .ots files, anchored in Bitcoin, verifiable with "ots verify"), and the scored run began at 07:01:59 UTC the same day.
Identifiers
Frozen commit: 3b7bcabf6c4ef71adc95e7e5546d3073612c2e88 (tag prereg-v1-frozen)
Binding commit, Amendment 11: ed4655f9cc5dbf0bd7ea90b5929a72794d5a9495 (tag prereg-v1-binding)
SHA-256 of PREREGISTRATION.md: e19aecb696d780d16d05363b34d850b9a58282097cd84b0b207df7eb77146014
Repository: https://github.com/GautamTalksDev/jevbench
Files: PREREGISTRATION.md (the plan and all amendments); PREREG_CHANGELOG.md (every edit after the binding commit, classified; all clerical); preregistration.lock.json (dataset hashes); jevbench-prereg-v1-frozen.zip (repository snapshot at the frozen commit); SHA256SUMS.txt; two OpenTimestamps proofs.
Notes. ChaosNLI sentence text is not included (CC BY-NC 4.0); it is fetched from the original source by scripts/fetch_chaosnli.py. runs/offline_fixture in the snapshot is a synthetic demonstration fixture added on 2026-09-23, before any access to Jev; its model field is a schema placeholder and it contains no real API responses. Code was written with AI assistance (Cursor and Claude); the author reviewed the design, code changes and all hashes.
Files
jevbench-prereg-v1-frozen.zip
Files
(2.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:c4791e72a3fc0c20591173785d431439
|
2.7 MB | Preview Download |
|
md5:145789accb8b20ba35f0af0381be2b83
|
4.1 kB | Preview Download |
|
md5:8a4d437c3173759a9833696a5da69178
|
563 Bytes | Preview Download |
|
md5:0a576f189ec1364b6241fa01a2ecbf75
|
29.2 kB | Preview Download |
|
md5:695712b6c95c3b723201916866f22902
|
3.9 kB | Download |
|
md5:053ab7d59151b5f6f90aa32e39b36d98
|
359 Bytes | Preview Download |
|
md5:d408ec0fe0455b3f383d14db636dd746
|
3.9 kB | Download |
Additional details
Related works
- Is referenced by
- Preprint: 10.5281/zenodo.22971492 (DOI)
- Is supplemented by
- Software: https://github.com/GautamTalksDev/jevbench (URL)
- References
- Conference paper: 10.18653/v1/2020.emnlp-main.734 (DOI)
Software
- Repository URL
- https://github.com/GautamTalksDev/jevbench
- Programming language
- Python
- Development Status
- Active