Published April 26, 2026
| Version v1
Preprint
Open
Training Neurosymbolic Beings: A Hybrid Symbolic-LLM Pipeline for Knowledge Extraction and Provenance Attribution
Description
This paper presents a hybrid training methodology for neurosymbolic AI beings achieving 95-100% provenance attribution while requiring 100-1000x less time than LLM fine-tuning. A 10-step pipeline extracts structured knowledge from domain documents and crystallizes it into inspectable RDF graphs in 8-16 minutes on CPU hardware. Through three domain-specific beings (software developer, AI ethicist, software architect), the authors validate a layered corpus strategy and a Q&A augmentation pattern where Q&A data validates and strengthens prose-derived understanding rather than replacing it, achieving 95-100% coverage on behavioral scenarios with full provenance attribution.
Files
PAPER-102-Training-Neurosymbolic-Beings.pdf
Files
(1.5 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:4d71cf34e0db210af9e8264468ede547
|
1.5 MB | Preview Download |