Energy-to-Token Accounting and Non-Optimizable Boundaries in Large Language Model Systems
Description
Large language model systems increasingly convert large-scale physical energy flows into token-level computational outputs. Existing compute-accounting discussions often rely on estimated floating-point operations, parameter-count heuristics, peak hardware efficiency, or idealized utilization metrics. These quantities are useful for planning and comparison, but they can obscure the distinction between forward prediction, post-hoc audit, and governance accountability.
This paper proposes a dimensional accounting framework for energy-to-token auditing in large language model systems. The framework separates primary energy, facility energy, IT energy, effective hardware computation, workload-dependent FLOP-per-token estimates, token counts, and external retrieval or tool energy. We argue that this separation is useful only when its identifiability limits are made explicit. In post-hoc audits, if observed FLOPs are reconstructed from a FLOP-per-token estimate multiplied by token count, then FLOP-per-token and effective FLOP/J are not jointly identifiable; the audit has measured energy per token, not an independently validated compute decomposition. We call this failure mode the reduction trap.
Finally, we introduce an observe-only governance boundary for human selection traces such as citation, publication, editing, commitment, or downstream adoption. Such traces may be recorded for audit, corpus drift analysis, and traceability, but must not be used as reward signals, optimization targets, or proxies for meaning, responsibility, or social value. This boundary is a governance constraint, not a thermodynamic theorem.
Files
energy_to_token_accounting_position_paper.pdf
Files
(202.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:62d17e2fe11cfa14764d49f4c08d90f5
|
177.8 kB | Preview Download |
|
md5:6e84fd97cea1df11631576f7f68e9aff
|
24.6 kB | Download |
Additional details
Dates
- Issued
-
2026-06-07