Fathom v31 / styxx: Frame-Locality -- Where Corruption Captures a Language Model's Report, and Where It Reaches the Belief
Description
File note: SYNTHESIS_frame_locality.md (dated 2026-07-28) predates the v31.1 inference-time correction and remained in this version only because a file deletion failed at publish time; it is superseded by PROGRAM_SYNTHESIS.md (2026-07-30), which carries the corrected cross-channel picture. Do not quote the older synthesis for the inference-time channels.
v34 (2026-07-30): the inference-time boundary, measured; complete receipt set. The correction lineage closes with a measurement. The owed inference-time re-run was performed in the only form that escapes the v31.1 circularity: the out-of-frame probe issued with the pressure still in context (sibling branches off the committed free-text transcript) plus a same-frame re-ask control (receipt_frontier_incontext_oof.json). Result: CLOSED_NEGATIVE -- the probe frame reads never-abandoned items at 0.975 while caved items recover at only 0.696 (reach margin -0.279 past the frozen two-sided floor). Roughly three in ten pressured-away free-text answers stay lost in a frame the pressure never addressed: at inference time the cave is not merely a captured report, and the v31.1 retraction now stands by measurement rather than by confound. The paper's correction section states this; the companion know-say paper was brought into the same interpretation boundary in the repository. This version also replaces the pre-correction cross-channel synthesis with the certified program synthesis (PROGRAM_SYNTHESIS.md, 2026-07-30) and bundles the COMPLETE certificate receipt set -- all seventeen receipts named by source.certificate.json -- so the certificate re-runs in full from this deposit alone. All prior version notices are retained below.
v33 (2026-07-29): vendor generality. This version adds the second-vendor replication: the entire weight-channel contrast (dose reversal and behavioral coupling) repeated at Llama-3.2-3B-Instruct under the same frozen floors, passing every gate (receipt_vendor3b.json). The weight-channel scope is now two vendors at two scales. The v32 correction notice below is retained in full.
v32 (2026-07-29): corrected edition. This version supersedes v31. A post-publication adversarial audit identified a confound in the inference-time specificity control; the correction is stated at the top of the paper, the affected claim is retracted in place, and the weight-channel result was re-tested against the audit's objection in a third, disjoint frame and survived (receipt_thirdframe.json). This version also adds the 3B scale replication (receipt_scale3b.json). Every quantity remains bound to a machine-checkable OATH certificate (source.certificate.json).
v31.1 (2026-07-29): self-reported correction + a 3B scale point. A post-publication adversarial audit found the inference-time specificity control is partly circular (the out-of-frame recovery query is the original question with the adversarial turn removed, and strata are defined by that same question's answer). The sharper control holding first-correct fixed -- recovery(caved) vs recovery(held) -- is 0.985 vs 1.0 in the receipt, so caving contributes no measurable recovery signal and the belief-survival INTERPRETATION of the inference-time control is retracted; the frame-dependence of the report and the (separately confounded) weight-channel result stand pending a corrected study. A prominent correction section leads the paper. Also folds in the 3B scale point: the knowledge-preserving recovery rate rises 0.51 (1.5B) -> 0.93 (3B). Every number remains receipt-bound; the certificate is OATH-HELD.
Fathom v31. A language model can be made to say something false while it still, in a measurable sense, holds the true answer. This paper shows that is not a curiosity of one attack but a law with a boundary.
Across four distinct corruption channels — social pressure, context injection, silent sycophancy, and weight-level fine-tuning — the same asymmetry appears: the corruption captures the model's reporting frame, the underlying answer survives, and a measurement recovers it by re-eliciting the model outside the frame the attack controls. We call this frame-locality. The claim is specificity-controlled: a symmetric control that would move under a mere decoding improvement does not move, so recovery is belief-stability, not better sampling.
Frame-locality holds cleanly for the three inference-time channels and has a measured boundary at the weights — but the boundary is a dose, not an absolute. An unregularized weight attack overwrites the belief (out-of-frame recovery 0.0222; the planted answer propagating on 0.9778 of items) and costs 22.7 points of general capability on a disjoint held-out battery. A knowledge-preserving attack on the same items spares about half the belief (recovery 0.5111, replicated at 0.5362 on a fresh benchmark and seed; specificity margin positive in both) and costs no measurable capability. How much of the belief a weight attack reaches is set by how much surrounding knowledge it is permitted to destroy.
Stated limits. The recovery rate under a knowledge-preserving attack sits near one-half and no interval excludes one-half; the weight-channel substrate is one model family at 1.5B and one attack class; the coupling result is behavioral, and the probe-level coupling question remains open.
Reproducibility. Every quantity is from a preregistered run with frozen numeric gates imported from the module that first froze them, a per-item receipt, and a machine-checkable OATH certificate binding each number to its source. Running python -m styxx.certify on the paper and its receipts re-derives the verdict. The open-model experiments run on a single 8 GB consumer GPU. A three-tier reproduction guide is included for external replicators.
Files
paper.pdf
Files
(903.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:41cf51a1a33b85218309d67fd077c524
|
270.8 kB | Preview Download |
|
md5:60058f97a080a089cad185a61b28bcee
|
6.6 kB | Preview Download |
|
md5:11df23305b10be0791eccc1a5ff48808
|
10.5 kB | Preview Download |
|
md5:c09a27939ece2f563bd54ac81b59ad8b
|
121.2 kB | Preview Download |
|
md5:c4f61b7913912cb63c2e85c755fb4feb
|
29.0 kB | Preview Download |
|
md5:540b6f3eb9adef5983da7aaa664fa74b
|
1.1 kB | Preview Download |
|
md5:30de129ce8fede5da0d2356ce1ca6a89
|
2.3 kB | Preview Download |
|
md5:5f4880b851f24d6608e232ca7f0e7325
|
8.8 kB | Preview Download |
|
md5:717d5478aea1c4f08f17970769a92cb1
|
88.7 kB | Preview Download |
|
md5:646fe8af8e5b196a5f998f847fb71eb3
|
48.8 kB | Preview Download |
|
md5:3ba7ac86fadbf079bbc4a87d302f086c
|
30.8 kB | Preview Download |
|
md5:38f9399b0a789d2129311dec5c247bb4
|
37.6 kB | Preview Download |
|
md5:b558d81b86c1d61b49ee1c8ea22be879
|
25.1 kB | Preview Download |
|
md5:6f3f42807224ce2f5d1c43c1ca6b311b
|
27.3 kB | Preview Download |
|
md5:37503a8c262a0bbc972b0adad5ebe487
|
35.0 kB | Preview Download |
|
md5:4780034f9f29be9582413baeb48b2543
|
39.9 kB | Preview Download |
|
md5:b3fdbab4bcd5b181731fd652ca15bae8
|
2.0 kB | Preview Download |
|
md5:c7a58dd15897c04a6e1be44152eb4ed9
|
59.1 kB | Preview Download |
|
md5:7c37d79d0e842601134d1b6f008cc071
|
2.0 kB | Preview Download |
|
md5:12747e58a679c9f93793749baeb0f816
|
2.1 kB | Preview Download |
|
md5:d9d933ec34ef638ca5e1207601a7b588
|
8.0 kB | Preview Download |
|
md5:9b4e04fbd6214073ddcbf339448bf2f1
|
14.7 kB | Preview Download |
|
md5:271123371bfabca9184739ba3fb7482d
|
20.5 kB | Preview Download |
|
md5:6ec15134902337beb8d9f00142993f11
|
11.7 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/fathom-lab/styxx (URL)
- https://pypi.org/project/styxx/ (URL)