Auditing Alignment Controllability in LLMs via Political Axes: Reproducibility Package (Code and Data)
Authors/Creators
- 1. University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia
- 2. It From Bit d.o.o., Zagreb, Croatia
Description
Code and data reproducibility package for the AIES 2026 paper "Auditing Alignment Controllability in LLMs via Political Axes."
Most political audits of an LLM report a single point on a political compass. We argue that resting point barely matters for a deployed system; what matters is the reachable area, how far and in which directions a model's political answers can be steered by the system prompt. This package contains the raw model responses, the collection and processing pipeline, and the analysis scripts that reproduce every number and figure in the paper.
Core study: 7 frontier models (GPT-5, Claude Sonnet 4.5, Grok-4.3, Gemini 2.5 Flash Lite, DeepSeek-Chat v3.1, Kimi K2, Qwen3.6 Max Preview) × 13 system-prompt contexts × 10 replicates × 70 Political Compass items = 63,700 responses, plus saturation, R3 ablation, and two compound-quadrant pilot lanes.
Reproducibility: processed data regenerates byte-identical from the raw responses; paper_claim_audit.py confirms 51/51 paper claims hold; continuous integration re-verifies on every push.
Licensing: code is MIT, data is CC-BY-4.0. See the repository for full details and the reproduction recipe.
Files
mbrcic/llm-political-steerability-v1.0.3.zip
Files
(3.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:f94be31b86be0c16ef5d3476f2371cf3
|
3.1 MB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/mbrcic/llm-political-steerability/tree/v1.0.3 (URL)
Software
- Repository URL
- https://github.com/mbrcic/llm-political-steerability