Toten: A Knowledge-Based System For Structure-Preserving Representation Of Physical Quantities And Technical Notation In Brazilian Portuguese
Description
Artificial intelligence pipelines performing quantitative reasoning over technical text depend on input in which physical quantities, numbers, units, and symbolic expressions arrive intact; when these entities are fragmented at tokenization, the error propagates throughout the downstream system. Statistical tokenization by Byte-Pair Encoding, optimized for vocabulary compression, is semantically blind to these entities and fragments them into lexically arbitrary subwords — a problem aggravated in technical Brazilian Portuguese. We present TOTEN, a knowledge-based system that produces an input representation in which each technical entity is preserved as a whole, typed unit: rather than deriving vocabulary statistically, it classifies textual regions declaratively, governed by a formal ontology of engineering entities (OEE) that constitutes the engine of the system. The ontological core is formalized as the triple ⟨O, classify, {instτ }⟩—types, structural principles, composition relations, and preservable invariants; a classification function that maps raw text into typed regions; and an indexed family of instantiators that produces a self-descriptive structured representation. Integrity rests on deterministic coupling to three consolidated external authorities: Pint (dimensional), the Unicode Character Database (typographic), and RSLP (Portuguese morphology). We evaluate the system on four properties verifiable by construction—ontological atomicity, dimensional equivalence, typographic robustness, and numerical reconstruction — over an internally generated and physically validated benchmark (EngQuant, N = 800) and four external corpora in Brazilian Portuguese (N = 1 771 cases eligible for numerical reconstruction), additionally reporting detection recall. Compared to eight representative state-of-the-art systems, TOTEN achieves unit ontological atomicity in all contrasts and numerical reconstruction of 0.775 to 0.904 on external corpora, against 0.627–0.703 for the best baseline (Quantulum3); on the internal benchmark, 0.780 against 0.340. Differences in atomicity and reconstruction are statistically significant (McNemar with Holm correction). The Spearman rank correlation between internal and external corpus rankings confirms the concurrent validity of the control benchmark. In dimensional equivalence, the system shows statistical parity with Pint, the oracle from which it inherits dimensional authority. These results characterize TOTEN as a structurally faithful, auditable, and computationally inexpensive input layer for intelligent systems operating on technical knowledge, without dependency on generative models.
Files
toten_preprint_v2.pdf
Files
(1.3 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:1900c5f7d997b90ccd3d0c18b0ef7ae7
|
1.3 MB | Preview Download |