Visual Oracle Bench (v1.8.0) — Clean Closed-Weight Re-Dispatch, Three-Judge Synthetic-HTML Corpus, and Silent-Fabrication Correction Protocol
Authors/Creators
Description
What v1.8.0 adds. A clean re-dispatch of both closed-weight judges (Claude Sonnet 4.5, OpenAI gpt-5-codex) under the patched wrappers with the fail-fast gate armed, superseding the contaminated original run as the basis for all Section 4 numbers. Both judges now have 553 clean judgments over a shared 553-pair subset (400 defect + 153 identity controls) across all eight profiles. Headline clean results: recall 70.5% (codex), 47.8% (claude), 7.5% (qwen); specificity 100% for all; paired full-corpus 96 codex-only vs 5 claude-only detections, exact McNemar p = 6.6e-23 (the earlier draft’s 25/0 strict-dominance claim was a contamination artefact and is corrected); defect-only Cohen’s kappa 0.504. New in this build: a controlled fault-injection study (Section 5.6) — analysis/fault_injection_study.py and analysis/results/fault_injection_matrix_2026-07-11.json — mapping eight wrapper fault classes against four integrity signals. A defect-only corpus with no added instruments misses five of the eight faults; identity controls and known-defect anchors are polarity duals, so a dual-polarity control set is required to cover the detection surface. See the bundled visual-oracle-bench-v1.8.0-redispatch.tar.gz (re-dispatch parquets, redispatch_reconcile.py, clean per-category data, regenerated Figure 1, and resume manifests) with README and SHA256SUMS. Every number regenerates from the parquets with no LLM call.
Replication package for the manuscript Silent Fabrication in LLM-as-Judge Pipelines: Identity Controls as Integrity Instruments, with Evidence from a 600-Pair Visual-Regression Case Study.
What this version adds over v1.6.1 (DOI 10.5281/zenodo.20645248). The v1.6.1 record was minted on 2026-06-11 and contains the closed-weight materials only. The open-weight dispatch was executed on 2026-07-07/08 and therefore could not appear in it. This version adds:
analysis/qwen25vl_dispatch.py— Qwen 2.5 VL 7B open-weight dispatcher (Ollama)results/judgments_qwen25vl7b_*.parquet— 600 open-weight judgmentsanalysis/results/three_judge_analysis_2026-07-07.json— three-judge confusion matrices, pairwise and three-way agreementanalysis/results/gwet_ac1_2026-07-07.json— Gwet AC1 with bootstrap CIs
The Qwen rows of Tables 1–3, the three-way Fleiss' kappa, and every Gwet AC1 value in the manuscript regenerate from this record. The closed-weight parquets, including the retained pre-correction originals that document the silent-fabrication contamination (817 of 1,400 dispatched judgments, 58.4%), are carried forward from v1.6.1.
Pre-registration. OSF 10.17605/OSF.IO/CSKUY. The earlier registration 10.17605/OSF.IO/NKD6J failed OSF archival on 2026-06-07 and was never finalized; it is superseded and should not be cited.
Source: github.com/SuneetMalhotra/visual-oracle-bench @ v1.7.0-stvr. Code MIT, data CC-BY-4.0.
Files
Files
(2.0 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:441277a572458ec90d8da7aa7e622b2b
|
1.7 MB | Download |
|
md5:27061ad89f949215b401dc82332e54b4
|
234.4 kB | Download |
Additional details
Related works
- Is new version of
- 10.5281/zenodo.20645248 (DOI)
- Is supplement to
- 10.17605/OSF.IO/CSKUY (DOI)
- Is version of
- 10.5281/zenodo.20620870 (DOI)