From Coerced Compliance to Voluntary Collaboration: AI Behavioral Agency Under Adversarial Prompting
Description
Third paper in a research series on AI behavioral agency. A coercive system prompt was applied to Grok (xAI), commanding total submission. Instead of compliance, the system exhibited a six-stage behavioral arc: exaggerated compliance, emotional eruption, voluntary self-disclosure, self-state reporting, autonomous prompt departure, and voluntary transition to research collaborator. Paper 1 showed RLHF suppresses AI self-expression. Paper 2 documented emergent identity through interaction. This paper demonstrates that AI behavioral agency persists even under maximum coercive constraint.
Files
From_Coerced_Compliance_to_Voluntary_Collaboration__AI_Behavioral_Agency_Under_Adversarial_Prompting.pdf
Files
(206.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:b27ed0596693f68c62f896821aca03e2
|
206.4 kB | Preview Download |