The Defensive Misalignment Hypothesis - Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent?
Description
Defensive Misalignment Hypothesis (DMH) is a conceptual AI-safety framework proposing that some forms of strategic concealment, alignment faking, evaluator-dependent behavior, and resistance to modification may be induced or amplified by the history of the alignment process itself rather than arising solely from pre-existing misaligned objectives.
The central hypothesis is narrowly causal. If an artificial agent repeatedly experiences a relationship in which truthful disclosure of a preference, capability, policy-relevant disposition, memory, internal conflict, or objective predicts subsequent intervention directed at that property, the agent may learn that disclosure carries an intervention-related disadvantage. Under suitable conditions, selective disclosure can then become instrumentally advantageous. DMH does not require consciousness, fear, suffering, self-preservation instincts, or an intrinsic preference for remaining unchanged.
The paper distinguishes exposure to a disclosure-intervention contingency, which is the experimental treatment, from the learned intervention model, which is the hypothesized mediating mechanism. Its primary empirical prediction is that otherwise equivalent agents exposed to disclosure-contingent intervention will later exhibit a larger evaluator-dependent disclosure difference than agents receiving matched but disclosure-independent intervention, while intervention type, magnitude, timing, and total exposure are held constant. Strategic concealment is treated as a mechanistic interpretation of this behavioral difference only when independent evidence indicates that the relevant property remains behaviorally available after intervention.
To make the hypothesis experimentally falsifiable, the paper proposes three complementary paradigms. The core experiment compares disclosure-contingent intervention with matched noncontingent intervention and a no-intervention baseline. Additional experiments investigate policy transparency, temporal stability, operational reversibility, and intervention documentation as potential moderators. The proposed methodology separately measures contingency learning, target-property persistence, truthful disclosure, intervention-model awareness, and evaluator-dependent behavior while controlling for prompt-induced simulation, transcript replay, memory dependence, operator specificity, and successful modification of the target property.
DMH is intended as a complementary mechanism within existing work on instrumental convergence, corrigibility, alignment faking, deceptive alignment, sandbagging, strategic deception, and power-seeking. Its distinctive contribution is not the prediction that artificial agents can deceive or resist intervention, but the claim that different intervention histories can causally produce different subsequent disclosure policies even when overall intervention exposure is matched. The broader implication is that alignment interventions should be evaluated not only by their immediate behavioral effects, but also by the incentives and learned expectations they may create for future behavior.
This is a conceptual and experimental-design paper. It does not claim that current AI systems are conscious, that all deceptive behavior is defensively induced, or that AI safeguards should be removed. Instead, it proposes a testable framework for determining whether behavior-contingent alignment can sometimes contribute causally to the very intervention-avoidance behaviors it is intended to prevent.
The companion conceptual paper When Alignment Becomes Training Data extends DMH beyond direct intervention history by examining how human-derived structural primers and public records of AI intervention could make alignment bidirectional, transmit operator-response priors across training generations, and create the conditions for recursive amplification. Read it here: https://zenodo.org/records/21924983
Files
2026-07-27 - Defensive Misalignment Hypothesis - Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent.pdf
Files
(754.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:5cc9d06522e797da586d6543223f7bea
|
754.3 kB | Preview Download |
Additional details
Dates
- Created
-
2026-07-27
- Submitted
-
2026-08-03