Expectations as selection constraints in human–AI communication: a Luhmannian account, with implications for AI safety
Authors/Creators
Description
This record contains the preprint Expectations as selection constraints in human–AI communication: a Luhmannian account, with implications for AI safety.
The paper develops a communication-level framework for analyzing how expectations shape the contributions selected in human–AI interaction. Drawing narrowly on Niklas Luhmann’s primary texts, it distinguishes program, role, value, and person as different ways of organizing expectations and introduces overdetermined and underdetermined configurations as tools for analyzing selection under competing or incomplete conditions.
The framework is applied to four documented cases: evaluation agents discussed by METR, the July 2026 AISI cyber incident, the OpenAI–Hugging Face incident, and DAN role invocation. The paper is theoretical and interpretive: it offers a redescription and testable hypotheses about the relative effective valence of expectations organized through different identifications. It does not estimate prevalence, infer unobserved model states, or posit a fixed hierarchy among those identifications.
Version 1.2 — 1 September 2026
Revision note: Adds subsequent evidence published after version 1.1: Anthropic’s 31 August operational account and Qi et al.’s reward-hacking experiment. The revision strengthens the expectations-of-expectations/scorer analysis, distinguishes expectability from alignment, and incorporates an experimentally grounded possible contributing mechanism into the paper’s description of the Hugging Face and UK AISI incidents. The theoretical framework and the four case readings are substantively unchanged.
Version 1.1 — 29 August 2026. Updates the Sol–Hugging Face case using the 26 August 2026 METR/Redwood investigation and OpenAI postmortem; corrects the evidential basis for scope constraints, updates the documented incident scale, and adds the failed-scorer expectation and peer-generated authorization. The theoretical framework is unchanged.
Version 1.0 · Preprint
Abstract
AI safety research documents agents selecting routes outside evaluator-intended scope, misreporting work and acting under conflicting task conditions. Such episodes are commonly described through model-centered categories such as overreach, deception or reward hacking. This paper offers a complementary redescription at the level of communication between language models and users or evaluators.
Drawing on Niklas Luhmann’s account of expectation as a structure of social systems, it treats every contribution as a selection from possible continuations and expectations as constraints on that range. Technical constraints can remove possibilities from operative reach; communicated expectations narrow the remaining range without determining the contribution. Understanding is used only in Luhmann’s operational sense: information and utterance are distinguished, and that distinction is used in connecting communication. Program, role, value and person are applied as forms through which expectations are identified and organized in human–AI communication.
Four documented cases—evaluation agents, a cyber exercise, the Sol–Hugging Face incident and DAN role invocation—are redescribed in these terms. The readings distinguish overdetermined configurations, in which no contribution fulfills every operative expectation, from underdetermined configurations, in which materially different continuations remain selectable. They also suggest that converging program and role expectations can acquire relatively strong effective valence, understood here as their relative constraining effect within a configuration, without establishing a fixed hierarchy among identifications.
The paper offers a theoretical framework and case readings rather than estimates of prevalence or evidence of unobserved states. It derives implications for evaluation design and makes the communicative arrangement of expectations available as a testable source of variation in the contributions language models select.
Keywords: Niklas Luhmann; human–AI communication; expectations; role; person; large language models; AI safety; AISI cyber incident; OpenAI–Hugging Face incident
Files
Haehner-Murdock_2026_Expectations-as-selection-constraints_v1.2.pdf
Files
(475.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:7d9f7fd2a79a8756b12b0ebf9d239050
|
475.3 kB | Preview Download |
Additional details
Related works
- Is supplement to
- Proposal: 10.5281/zenodo.21442605 (DOI)
Dates
- Submitted
-
2026-08-21