Published May 20, 2026 | Version v1

PolyglotQL: A Pipeline for Multilingual Text-to-SPARQL Dataset Generation.

  • 1. ROR icon Technische Universität Berlin
  • 2. Deutsches Forschungszentrum für Künstliche Intelligenz GmbH
  • 3. ROR icon German Research Centre for Artificial Intelligence

Description

We present PolyglotQL, an open-source ETL (Extract, Transform, Load) pipeline for systematically creating
multilingual text-to-SPARQL datasets, along with an accompanying framework for evaluating text-to-SPARQL
generation models. PolyglotQL provides an extensible and modular architecture that aggregates, normalizes, and
augments heterogeneous question–SPARQL pairs from established text-to-SPARQL datasets. With this pipeline, we
automatically construct a bilingual English–German dataset featuring contextualized entity and relationship mappings
as well as automatically translated and aligned question pairs. We also conduct an empirical evaluation using two
multilingual open large language models under two distinct contextualization settings. The results show consistent
performance improvements when explicit grounding information is provided, highlighting the benefits of structured
context in multilingual semantic parsing.

Files

PolyglotQL_A Pipeline for Multilingual Text-to-SPARQL Dataset Generation.pdf

Additional details

Funding

European Commission
HIVEMIND - Human-centred collaboratIVE MultI-ageNt framework for accelerating software Development and maintenance 101189745