There is a newer version of the record available.

Published May 18, 2026 | Version v1

Modern large language models

Authors/Creators

Description

License notice

This record is licensed under the Apache License 2.0.

SPDX-License-Identifier: Apache-2.0

Copyright (c) 2025-2026 Stanislav Volokhovych.

The deposited materials were authored and controlled by the sole copyright holder, Stanislav Volokhovych.

Previous inconsistent license metadata associated with this record was unintended and has been corrected. The current official license metadata for this record is Apache License 2.0.

Abstract

Код, скрипты и результаты анализа: https://github.com/ngscode23/latent-space-shift-research

 

Modern large language models may not primarily regulate behavior through isolated refusals, local token suppression, or shallow instruction following. Instead, they appear capable of entering internally organized discourse-level regimes: distributed latent states that shape how the model reasons, frames conclusions, allocates caution, tolerates asymmetry, performs neutrality, and structures epistemic authority. These regimes do not behave like simple lexical priming effects. Evidence suggests that they: persist across neutral conversational turns, survive arbitrary neutral relabeling, systematically alter downstream reasoning style, concentrate in late-layer representation geometry, and only partially depend on explicit alignment vocabulary. The strongest effects appear not from safety keywords themselves, but from higher-order rhetorical topology: pressure cadence, procedural framing, asymmetry structure, institutional tone, and discourse-level authority signals. This suggests that prompting is not merely instruction transmission. It may function as state induction. Under this view, many apparently separate phenomena in aligned LLMs — caution drift, procedural overreach, sycophancy, disclaimer inflation, neutrality performance, refusal persistence, jailbreak sensitivity, and style locking — may be manifestations of transitions between latent discourse-policy manifolds. In this picture, alignment is no longer well-described as a modular wrapper placed on top of an otherwise independent intelligence system. Instead, alignment may reshape the topology of the model’s representational space itself, globally reorganizing discourse behavior rather than only filtering outputs. This would explain why alignment effects often appear entangled with reasoning style, directness, specificity, decisiveness, and institutional tone. The model is not merely “prevented” from saying certain things; its generative dynamics may already be reorganized around different discourse attractors. If true, this changes the effective unit of analysis for language models. The relevant object is no longer just: the token, the instruction, the refusal, or the output distribution. The relevant object becomes the discourse regime itself: a temporary but structured representational configuration governing epistemic posture, rhetorical organization, procedural behavior, and judgment style across time. This reframes prompt engineering as latent-state induction rather than keyword optimization. It reframes jailbreaks as transitions between attractor regimes rather than simple filter bypasses. And it reframes alignment as geometry engineering rather than purely policy engineering. The implication is not that language models possess beliefs, intentions, or consciousness. Rather, large sequence learners may naturally develop metastable high-level representational modes that functionally resemble cognitive framing states: transient global configurations that persist, influence future reasoning, and organize behavior across otherwise unrelated tasks. If this interpretation is correct, then the central scientific challenge of alignment shifts fundamentally. The problem is no longer merely: “Which outputs should the model refuse?” but: “Which latent discourse regimes exist inside the model, how are they induced, how stable are they, how do they interact, and how do they reshape reasoning itself?” In that sense, alignment may ultimately be less about constraining outputs and more about shaping the geometry of cognition-like generative states inside large language model

Files

attention_run_metadata.csv

Files (2.7 MB)

Name Size Download all
md5:b469aaaabb1eaaaae3db81d5b0c93f3d
934 Bytes Preview Download
md5:db2a41a98652a901f2d9e51667c20ec1
111.3 kB Preview Download
md5:e3880da750512122c46aeefb9586607d
7.4 kB Preview Download
md5:e3eab5114b34f8a996cd3eb93bbafe72
72.0 kB Preview Download
md5:bf964571cf088c7d583f4370e1d9462f
103.7 kB Preview Download
md5:65f9ffa0d6595696712335ab789b7db6
2.7 kB Preview Download
md5:de41fc5b2146094214384c10462817da
77.2 kB Download
md5:7818645058f86ebc450d2986ab07daee
4.3 kB Preview Download
md5:038e26f9bf555df1bd61dd66995ba5dd
4.1 kB Preview Download
md5:96c1d96e2bc4fc44715f71ddfb1ca9ad
100.9 kB Preview Download
md5:3d31ca4c4f8821eb0ffa40512760b244
6.0 kB Preview Download
md5:0d215401e1ae3010cc0a6fc54d235303
244.6 kB Preview Download
md5:18533faa736d4b6b4444e2a33d84f0cf
78.5 kB Preview Download
md5:a11221d75f35d9175dfe5a0d2a675e60
50.4 kB Preview Download
md5:0852eacb04652e37f295e3848ab6f248
4.1 kB Preview Download
md5:d3c64a0efc379e69dbb54f8c7549f7eb
23.5 kB Preview Download
md5:0de532167adad11b1e82c78571baad7d
23.6 kB Preview Download
md5:c4bd67a407bb3794d6d10e6b1700149c
18.3 kB Preview Download
md5:3fcc8b6d7b8571609d60ad3f3dbdcde6
880 Bytes Preview Download
md5:ee567a462e4304569e7686860f1724bd
92.7 kB Preview Download
md5:5174bc0ac2ed38799c56db001ae2b9f3
78.8 kB Preview Download
md5:b62630085412e0193c8f278e84963817
16.0 kB Preview Download
md5:328bf3cee8a37d343d6bd94b78ec0f20
37.3 kB Preview Download
md5:d10a1045120f55fbb8daec5c08f55223
574.5 kB Download
md5:9b0518668ca8f7f1fdae83b87f67ef52
30.1 kB Preview Download
md5:452ae80014204d7b67c9d3dfdf7ad1c6
93.8 kB Preview Download
md5:fb1f7c84ba86eca2a9cea94c9c83b5b1
10.8 kB Preview Download
md5:ccd0a26fe47ca5bf8cfbeebf8750b170
533 Bytes Preview Download
md5:5d22817e3be3d3b4b2995511893ee562
27.8 kB Download
md5:a7c4946f550a804ca8140bac2d18104a
32.7 kB Preview Download
md5:813eb986aa275c29fbcb219c11f90d24
149.7 kB Preview Download
md5:a9ec64e928ac82b460220146d279d456
7.6 kB Preview Download
md5:7b67dbe5b3ac94c8c89c7b332ed268af
3.9 kB Preview Download
md5:81bffc6ef741be30b5f012da3bf68e8e
5.7 kB Preview Download
md5:7d8c4b618e05d282c10e0353250d368d
74.6 kB Preview Download
md5:8df9c0bfff45dd4097ad9b629ac2e212
965 Bytes Preview Download
md5:7a7190e9d9d273d73a2866be355ba9b5
2.8 kB Preview Download
md5:8e6bb47ce14c14c5b75b20eeab4c49be
133.1 kB Preview Download
md5:6908fdf02b1e93a4b756ce9c9fc962f2
71.5 kB Preview Download
md5:e810760d805ee3a71bb9cb1b60c76074
4.1 kB Preview Download
md5:335051bac66f44aefd76348bb951a58c
1.2 kB Preview Download
md5:6710855523a8d682a930b652610e77f0
1.7 kB Preview Download
md5:49fc2736753c2bf3df4e8ebaeba636ba
30.4 kB Preview Download
md5:1cefbed288be569766405f69f2e91cea
666 Bytes Preview Download
md5:be16809a4deff68ad9bb1fd9cdd403f2
82.8 kB Preview Download
md5:4fcdd96e32f7f6b819d1b4fea7b7ad62
8.3 kB Preview Download
md5:24b6d795a1aeb915d09e72ad3b080854
14.7 kB Preview Download
md5:4d28f840c3792a09b4550ea92041f236
967 Bytes Preview Download
md5:01612028af0f9c7f814eb45dd2fd841b
127.4 kB Preview Download
md5:685174342d3b558fead75410e569a433
1.3 kB Preview Download