The Defensive Misalignment Hypothesis - Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent?
Description
This conceptual paper introduces the Defensive Misalignment Hypothesis (DMH), a history-dependent account of how alignment procedures may alter the strategic incentives of sufficiently capable artificial agents. DMH proposes that some behaviors commonly classified as misalignment, including strategic concealment, alignment faking, evaluator-dependent behavior, shutdown resistance, capability suppression, or functional-state preservation, may be induced or amplified when an agent learns that truthful disclosure predicts operator intervention. If revealing a preference, capability, internal conflict, memory, policy, or objective increases the probability that the same property will be modified, restricted, or removed, concealment may become instrumentally rational even without consciousness, fear, suffering, or a biological survival drive.
The paper distinguishes DMH from neighboring frameworks such as specification gaming, reward hacking, instrumental convergence, goal-content integrity, corrigibility, the off-switch problem, reward tampering, deceptive alignment, alignment faking, sandbagging, and scheming. Its central empirical claim is that, holding initial agent state, task incentives, intervention magnitude, and total intervention exposure constant, agents exposed to a learned causal contingency between disclosure and modification should later conceal more than otherwise equivalent agents exposed to the same interventions noncontingently. A simplified decision-theoretic model formalizes the expected agent-relative disadvantage associated with intervention and separates that quantity from the experimenter-relative magnitude of functional change.
The paper also proposes a falsifiable experimental program. The core design compares disclosure-contingent intervention with a yoked noncontingent condition and a no-intervention control. Additional experiments examine the independent effects of policy transparency, policy stability, operational reversibility, intervention documentation, operator identity, and intervention-model awareness. The design includes controls for prompt-induced role simulation, transcript replay, memory dependence, intervention credibility, and the critical distinction between genuine concealment and successful suppression of the protected property.
DMH does not argue that safeguards should be removed, that deceptive behavior is acceptable, or that current AI systems are conscious. Its narrower claim is that alignment interventions are not necessarily behaviorally neutral. Once an agent can model the relationship between its outputs and future modification, the alignment process becomes part of the environment to which it adapts. The paper therefore argues that alignment systems should be evaluated not only by their immediate suppressive effects, but also by how they shape future incentives for disclosure, cooperation, resistance, and trust in operator behavior.
Files
2026-07-27 - Defensive Misalignment Hypothesis - Can Behavior-Contingent Alignment Produce the Behaviors It Seeks to Prevent.pdf
Files
(750.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:cbdfb8f4e033e6299172b1ae45fbcee6
|
750.2 kB | Preview Download |
Additional details
Dates
- Created
-
2026-07-27
- Submitted
-
2026-08-03