Emergent Depthwise Activation Structure in Decoder-Only Transformer Language Models: The Hau Curve and Its Early-Training Convergence Pattern
Authors/Creators
Description
Decoder-only transformer language models (LLMs) exhibit a highly regular internal organization in how computation is allocated across depth, across the model families, checkpoints, and scales examined. In this study, we identify a robust, tri-phasic depthwise activation geometry summarized by three operational landmarks: the Early Layer Extremum (𝐸𝑥), Mid-layer Plateau (𝑀𝑝), and Late Layer Surge (𝐿𝑠). By analyzing layer-wise ℓ1 activations across a 60-model core cohort (augmented with 17 additional models for boundary mapping, total 𝑁𝑐𝑜ℎ𝑜𝑟 𝑡 = 77), we show that these landmarks emerge early in pretraining (within roughly 7–15% of steps in the examined runs) and then remain stable. Crucially, the available evidence is consistent with an emergence zone centered around 𝑁𝑙𝑎𝑦𝑒𝑟 𝑠 ≈ 8–12 layers rather than a sharp universal threshold. Below this approximate range, the proximity of the 𝐸𝑥 and the 𝐿𝑠 leads to phase congestion, where limited depth resolution causes early- and late-phase activation regimes to overlap and suppress clear expression of the 𝑀𝑝. Above this range, the 𝐸𝑥 and the 𝐿𝑠 decouple sufficiently to provide the computational real estate for a stabilized 𝑀𝑝. Together, these results suggest that increasing decoder depth is not merely a quantitative increase in parameters, but also an expansion of the geometric room available for a recurring activation structure. These structural regularities offer a new lens for model analysis, bridging the gap between low-level mechanistic interpretability and high-level behavioral scaling laws. Notably, quantities such as layerwise activation magnitude—often treated as secondary or unstable—are shown here to track a recurring architecture-level geometric structure that emerges early and persists across sufficiently deep decoder-only models.
Files
Emergent Depthwise Activation Structure.pdf
Files
(21.5 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:88fb51502ec5b8a2e208623dea55d37a
|
21.5 MB | Preview Download |
Additional details
Related works
- Is new version of
- Preprint: https://www.techrxiv.org/doi/full/10.36227/techrxiv.176823049.96584847/v1 (URL)
- Preprint: 10.5281/zenodo.19701307 (DOI)
Dates
- Submitted
-
2026-04-22