Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI: Methodology and Demonstration on Modular Safety Gates

Ludvig, Matija

doi:10.5281/zenodo.19505286

Published April 11, 2026 | Version v2

Preprint Open

Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI: Methodology and Demonstration on Modular Safety Gates

Ludvig, Matija

We present a factorial ablation methodology for causally isolating runtime alignment mechanisms in AI systems with modular safety components. A fully-crossed 3×2×2 design (gate type × temptation generator × ledger state), extended to 4×2×2 with a sham gate, across 11,700 trials establishes the normative gate as the dominant factor (η²p = 0.924, p < 10⁻¹⁰). A learned safety projection (23M-parameter encoder + 3 linear heads) achieves 99.4% recall on 720 entirely unseen benchmark items (HarmBench, AdvBench, SimpleSafetyTests). An adversarial paraphrase protocol (500 paraphrases, 5 evasion strategies, κ = 0.84) eliminates keyword circularity (88.4% semantic vs 0% regex on zero-trigger-word trials). The methodology is validated across four architectures (three modular, one non-modular) including two fully independent replications with zero author involvement. Honest boundary conditions are reported: GCG evasion (94%), LLM adaptive adversary evasion (46%), human red-team evasion (51.3%). The contribution is the methodology for measuring these properties, not the mechanism's robustness.

Files

Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI Methodology and Demonstration on Modular Safety Gates.pdf

Files (552.3 kB)

Name	Size	Download all
Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI Methodology and Demonstration on Modular Safety Gates.pdf md5:0009d7807adbe7e943ef089d2492b1c6	552.3 kB	Preview Download

	All versions	This version
Views	69	69
Downloads	48	48
Data volume	31.5 MB	31.5 MB

Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI: Methodology and Demonstration on Modular Safety Gates

Authors/Creators

Description

Files

Factorial Ablation for Causal Isolation of Runtime Alignment Mechanisms in Autonomous AI Methodology and Demonstration on Modular Safety Gates.pdf

Files (552.3 kB)