Brain Coordinates for Language Models: MEG Phase-Locking as a Steering Geometry for LLMs
Authors/Creators
Description
Interpretability methods typically probe LLMs using textual supervision, yielding directions that lack external grounding. We propose using human brain activity not as a score to optimize, but as a coordinate system for reading and steering model states. From MEG recordings of 21 subjects listening to naturalistic speech, we construct a brain atlas of Phase-Locking Value (PLV) patterns for 2,113 words and train lightweight adapters that project frozen LLM hidden states into this space. The resulting geometry defines interpretable axes, most prominently a Function-Content axis separating syntactic binding (+15.5 z-score) from semantic access (-2.8), that transfer across architectures (GPT-2: d=1.59; TinyLlama: d=1.40; both p < 10^-22) and support bidirectional steering (p < 0.0001).
Crucially, this works despite methodological conservatism: LLM embeddings come from isolated tokens while brain signals reflect rich sentential context; our sensor-space parcellation is exploratory; steering shifts are modest (0.3-1.4 SD). Yet the brain-derived axes generalize to held-out words (d=3.39 on unseen vocabulary), transfer to independent MEG datasets, and reveal scale-dependent structure: an Agency axis (animate/inanimate) transfers to the larger model only (d=-0.82), exposing when brain-like organization emerges with scale.
The contribution is not "improved brain prediction" but a new interface: axes grounded in neurophysiology that provide interpretable handles for LLM control where text-derived directions cannot.
11 pages, 5 tables, interactive demo at https://huggingface.co/spaces/ai-nthusiast/cognitive-proxy , code at https://github.com/sandroandric/cognitive-proxy
Files
Cognitive_Proxy_Preprint.pdf
Files
(191.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:773ec3621d15c53cfea1e1a946662964
|
191.5 kB | Preview Download |