Published April 26, 2026 | Version v1

Training Neurosymbolic Beings: A Hybrid Symbolic-LLM Pipeline for Knowledge Extraction and Provenance Attribution

Authors/Creators

  • 1. Congruent Systems LLC

Description

This paper presents a hybrid training methodology for neurosymbolic AI beings achieving 95-100% provenance attribution while requiring 100-1000x less time than LLM fine-tuning. A 10-step pipeline extracts structured knowledge from domain documents and crystallizes it into inspectable RDF graphs in 8-16 minutes on CPU hardware. Through three domain-specific beings (software developer, AI ethicist, software architect), the authors validate a layered corpus strategy and a Q&A augmentation pattern where Q&A data validates and strengthens prose-derived understanding rather than replacing it, achieving 95-100% coverage on behavioral scenarios with full provenance attribution.

Files

PAPER-102-Training-Neurosymbolic-Beings.pdf

Files (1.5 MB)

Name Size Download all
md5:4d71cf34e0db210af9e8264468ede547
1.5 MB Preview Download