SAFE-AI Decision Diagnostics: Community-Conditional Representational-Adequacy Testing for Frozen-Weight LLMs
Authors/Creators
Description
SAFE-AI Decision Diagnostics (previously circulated as SAFE-AI Reversed) is the diagnostic companion to the SAFE-AI framework. The forward volume (Monograph I) treats the user as an active agent and a frozen language model as a fixed environment to be steered toward a target, showing how disciplined interventions move a model's output toward what a user wants. This volume reverses the inference. It asks the prior question the forward account leaves open by design: whether a given model has the geometric capacity to register and respond to a particular community's authorized target at all — and, when it does not, what kind of failure that is, and what evidence would license saying so.
The work's central contribution is the separation of two failure signatures that application-level evaluation routinely conflates. A flat failure (low sensitivity) is a model lacking internal geometric capacity along a community-relevant direction — an intervention pushing where the model cannot register it. A wrong-direction failure (low descent) is a model that is sensitive but steers toward the wrong target — an elicitation problem, not a capacity one. The framework operationalizes the distinction through a four-cell sensitivity × descent matrix built on a context-space pullback Fisher metric, Rayleigh sensitivity, a cotangent steerability reading, and an optimal-transport efficacy measure — with distributional displacement explicitly separated from target-directed descent — reported profile-first across the community's declared dimensions, with scalar adequacy scores as community-authorized summaries. A persistent flat-and-unhelpful signature is read as an operational verdict of representational inadequacy, carried by a sequential Persistent Flat-Failure certificate whose anytime-valid form remains sound under the escalation ladder's adaptive stopping.
A discipline runs throughout: the apparatus issues operational verdicts conditioned on a named intervention class, the auditor's access regime, and the target's version and rendering — never impossibility theorems. Because it operates through the text interface, its behavioral layer applies to any frozen-weight model by construction; a certificate-tier system (DD-0–DD-4) grades what the assembled evidence actually supports, from target-and-rendering audit through black-box behavioral findings to causal internal evidence, with structural representational certificates requiring the access they name. Version 3 adds a formal abstention state: Cell U, the verdict the apparatus returns when evidence is discordant or underidentified — because a diagnostic that cannot abstain cannot be fully trusted when it concludes. The escalation ladder is reread as a sequential identification strategy that partitions persistent failures into three remediation classes — prompt-side, retrieval-side, representational — while a compatible-cause set tracks which underlying causes remain; causal attribution is licensed only when the exclusion audits, including a six-channel SAG adequacy certificate, close that set to representation alone. When they do, remediation is specified as selective curvature allocation with a quantifiable, rank-indexed cost, and the boundary into persistent learned state is indexed by an adaptation profile that hands off to Monograph III's authorization machinery. The target against which all of this is measured is defined and authorized by the community, with data-sovereignty principles (CARE, OCAP, Te Mana Raraunga) entering as formal constraints rather than commentary, and refusal admitted as a legitimate terminal state.
The volume develops applications — a measurable science of prompt efficacy under an operational attribution rule, model selection and procurement, capability certification, closed-loop drift monitoring, and protocol-conformance certification — and closes with a minimal experimental program: conjectures with explicit falsifiers, the smallest study that would calibrate the instrument, and a validation map recording the calibration obligations of the Version 3 apparatus. Each section opens with a plain-language summary, forming a continuous, self-contained second reading alongside the formal development. The volume inherits the formal substrate of Monograph I by reference and interfaces forward to Monograph III.
This monograph has not been developed with, reviewed by, or endorsed by any Indigenous People, Nation, community, or governance body. It offers a proposed technical apparatus and defers to applicable community authority and protocol.
Files
SAFEAI_Decision_Diagnostics_Monograph_II_V3.pdf
Files
(1.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:905c631bece327cf8581af7f28f6fc9a
|
1.2 MB | Preview Download |
Additional details
Related works
- Is continued by
- Working paper: 10.5281/zenodo.21776916 (DOI)
- Is derived from
- Working paper: https://doi.org/10.5281/zenodo.20649477 (URL)
- Is new version of
- Working paper: 10.5281/zenodo.20938141 (DOI)