Published August 11, 2026 | Version V12.0

Persistent State Machine (PSM): A Formal Computational Paradigm and Secure Reference Architecture for Energy-Efficient Attention Computing (Version 12.0)

Authors/Creators

  • 1. Cosmos Administrative Scrivener Office & Independent Researcher

Description

Persistent State Machine (PSM): A Formal Computational Paradigm and Secure Reference Architecture for Energy-Efficient Attention Computing (Version 12.0)

ABSTRACT
This paper introduces the Persistent State Machine (PSM), a formal parallel computational paradigm designed to bypass the physical and energetic limitations of the Von Neumann architecture, specifically targeting the "memory wall" in modern Large Language Model (LLM) workloads. We establish the mathematical framework of the PSM and prove its computability, discrete error bounds, and its mapping to DSPACE(O(n)).

Furthermore, we present the Active State-machine Memory Architecture (ASMA) as a secure hardware reference implementation of the PSM, specifically optimized for LLM Key-Value (KV) cache attention computation. To address multi-tenant cloud vulnerabilities, we integrate a comprehensive hardware security suite, including a Triple Modular Redundancy (TMR) self-checking voter protected against synthesis optimization, a constant-time binary AND-tree comparator, active shielding, and Dual-Rail Precharge Logic (DRPL) to resist power analysis.

Finally, we demonstrate the physical viability of ASMA on a commercial AMD Alveo U200 FPGA accelerator card under Microsoft Azure NP10s instances. The synthesised and placed-and-routed reference implementation achieves full timing closure at 166–200 MHz, maintaining a worst-case negative slack of 2.548–3.182 ns. Under realistic LLM inference workloads emulated via WikiText-2, the proposed gating policy limits perplexity degradation to only +1.7867% at a 50% attention gating ratio (50% active cells). The self-checking TMR voter incurs a negligible logic area overhead of only +4.84% LUTs, while the dynamic power scales linearly with array activity (consuming only 21 mW at a 20% active ratio). The results demonstrate a secure, mathematically verified, and highly energy-efficient hardware paradigm for next-generation Transformer attention computing.

Japanese Patent Application No. 2026-184865 (Patent Pending)

Files

zenodo_IEEE_TC.pdf

Files (332.3 kB)

Name Size Download all
md5:98df5f983029643cecae88fb6c6aa6ab
332.3 kB Preview Download