A Self-Replicating Swarm of Tiny AIs: Phasor-Face Transformers (PrismFormer) with Arithmetic as Algebra and Bit-Exact, Mergeable Gradients
Authors/Creators
Description
I present PrismFormer, a transformer-shaped sequence model in which every token, word or number, is a "phasor face": a bundle of unit complex numbers composed by one algebraic operation, per-component complex multiplication ("binding"). Numbers are encoded so that binding performs arithmetic: addition on a linear-phase band, multiplication on a log-phase band, with results read back by correlation, so a computed value is decodable rather than merely reachable. This numeric representation is a direct application of the frequency-domain (phasor) form of Holographic Reduced Representations (Plate 1995; 2003) and of Random Fourier Features (Rahimi & Recht 2007); my contribution is to make it the trainable substrate of a transformer whose dense maps are replaced by parameter-lean "relation-banks". Because that substrate can be evaluated read-only during backpropagation, gradients accumulate into a detached buffer: a minibatch split into shards and summed is bit-for-bit identical to the serial result, which enables a full-replication training swarm, in which every node holds the whole model on a plain CPU, that stays mathematically coherent across mismatched hardware, unlike sharded systems that need matched accelerators.
I report results against a parameter-matched dense transformer at each setting. Arithmetic emerges from the codec with no training (binding two number faces decodes to their exact sum and product). Given the same worked, column-by-column problem, PrismFormer learns multi-digit addition, subtraction, and multiply/divide by a digit and generalises to unseen operand pairs where a matched transformer stays near zero; its working is legible one column at a time with no trained probe. On comparison and relational tasks it generalises better than a transformer that fits the same training data (65.9% vs 49.1% held-out), and it is a competitive character language model at matched size. A two-sided ablation shows that freezing the numeric identity does not help a single operation in isolation but wins under shared multi-task load. All results are at approximately 100k parameters; out-of-range extrapolation and large-scale convergence of the full colony remain open. Every quantitative result is produced by a script in the repository and reproducible with the command given in its section.
This record bundles a frozen source snapshot (repository tag paper-v1) that reproduces every number; the bundled code is under its own source-available LICENSE (see the LICENSE file in the archive).
Files
PrismFormer-repro.zip
Files
(483.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:1ec8b23e872d7d7fbd7e7b77a39bcb2b
|
59.9 kB | Preview Download |
|
md5:bf8fb623dff67799315d336ded473163
|
423.7 kB | Preview Download |