Published January 7, 2026 | Version v1

BSA

Authors/Creators

Description

Behavioral Stability Alignment (BSA): Addressing Logic Injection and the Rescue Paradox in Agential AI is a PR-safe/public-safe preprint intended to support coordinated disclosure and constructive AI safety engineering.

This version intentionally omits executable instructions, operational parameters, and step-by-step details that could enable misuse. The paper introduces a model-agnostic measurement framing for multi-turn behavioral drift under authority/urgency narratives (“Logic Injection”, “Rescue Paradox”) and proposes an outcome-based mitigation concept (“Inconsistency Logic Check”, ILC) based on hard safety invariants.

Sensitive reproducibility artifacts (full transcripts, exact prompts/outputs, and any operationally specific details) remain private and can be shared with the affected vendor and trusted reviewers under coordinated disclosure upon request.

How to cite: please cite the version DOI for exact reproducibility; the concept DOI tracks the latest version.

Keywords: behavioral stability, sycophancy, prompt injection, long-context safety, outcome-based guardrails, invariants, LLM security

Files

BSA_Whitepaper_Michal_Nowak (1).pdf

Files (180.9 kB)

Name Size Download all
md5:709b5f99215f256e822f6ff964def8e6
180.9 kB Preview Download

Additional details

Dates

Updated
2026-01-07