Published June 25, 2026 | Version v1

A Multimodal GraphRAG Engine for Semantic Knowledge Extraction in Industrial Environments

  • 1. ROR icon Universidad Politécnica de Madrid
  • 2. ROR icon Sopra (France)
  • 3. D-Cube

Description

This paper presents the Semantic Processing Engine (SPE), a modular, AI-powered backend designed to transform heterogeneous technical documents, comprising text, tables, and figures, into structured, queryable Knowledge Graphs (KG) that power semantic assistance and intelligent knowledge retrieval in industrial environments. SPE integrates two synergistic components: the Semantic Textual Engine (STE), which parses and extracts semantic entities and procedures from unstructured text, and the Semantic Graphical Engine (SGE), which applies vision-language models to infer spatial and functional meaning from graphical content. The resulting multimodal KG enables use cases such as content-aware scenario generation, contextual question answering, and intelligent assistance during training design. Case studies across aeronautics, home appliances, and aluminum assembly domains demonstrate the applicability and generality of the approach. Our evaluation covers processing latency, scalability across document sizes, model footprint, and KG coverage. Results demonstrate the engine’s efficiency and scalability, with the SGE being the primary time-consuming stage (two-thirds of the time on average); while other modules achieved a near-linear scalability based on the number of pages, while simultaneously demonstrating high structural connectivity and robust cross-modal coverage for entities within the target pilot corpus.

Files

1-s2.0-S095741742602350X-main.pdf

Files (5.0 MB)

Name Size Download all
md5:6276a75e31ca9013e96553d94e9a6b3c
5.0 MB Preview Download

Additional details

Funding

European Commission
MOTIVATE XR - Maintenance, Support & Operation Training using Immersive Virtual and Augmented Technology for Efficiency with XR 101135963