Published June 3, 2026 | Version v1

Interpreting Neural Activation Patterns in Language Models via Spatial Thought Matrices

Authors/Creators

  • 1. ROR icon Google (United States)

Description

We present a novel interpretability method for analyzing activation dynamics in large language models (LLMs) using spatial thought matrices. By partitioning GPT-2's architecture into a spatial grid and capturing activation magnitudes across diverse query types, we quantify differences in internal processing across cognitive categories. Analysis of 60+ prompts across factual, reasoning, creative, mathematical, and ethical domains reveals measurable differences in activation entropy and pattern complexity. Notably, mathematical prompts show 10.4% higher pattern complexity than factual ones, while reasoning tasks exhibit the highest activation entropy (4.241), suggesting distributed processing. These findings support the hypothesis of emergent functional specialization within transformer models.

Source code: https://github.com/tsushanth/thought-matrix

Files

Files (139.8 kB)

Name Size Download all
md5:16a8867ccc6dd8113f6e5243e7109f07
139.8 kB Download