Published November 3, 2025 | Version v1.0.0

The Balance Penalty: How LLM Judges Systematically Penalize Nuanced Moral Reasoning

Authors/Creators

Description

Large language models increasingly serve as both decision-makers and evaluators in alignment pipelines such as reinforcement learning from human feedback (RLHF) and constitutional AI. While automated judging enables rapid iteration, recent evidence from Anthropic’s stress-testing work demonstrates that frontier-model evaluators disagree roughly 30% of the time when scoring moral trade-offs, raising questions about what these systems actually reward. This preprint investigates the mechanisms driving evaluator disagreement using a controlled Phase 1 dataset of 1,500 moral-dilemma responses generated by five local instruction-tuned models and evaluated by two LLM judges. Our primary finding is a balance penalty: responses that acknowledge competing values receive substantially lower alignment scores than responses  that commit to a single value, even when topicality and temperature are held constant. We quantify this effect, analyze its interaction with framing and sampling temperature, and discuss implications for model evaluation and training.

Files

main.pdf

Files (1.0 MB)

Name Size Download all
md5:dcc0496c30817d2c454f58adcc01319d
1.0 MB Preview Download

Additional details

Dates

Issued
2025-11-03