Published July 31, 2026 | Version v1

GL-MedQuAD: A Curated English–Galician Medical Question Answering Dataset

Description

This data descriptor introduces GL-MedQuAD, the first domain-specific biomedical Question Answering (QA) benchmark for the Galician language. Comprising 2,100 parallel records derived from the MedQuAD corpus—specifically from GHR (N=1,467) and NIHSeniorHealth (N=633)—the dataset addresses the critical scarcity of specialized clinical NLP resources in low-resource linguistic scenarios.

The dataset was constructed through a hybrid, multi-stage generation-and-curation pipeline:

  • Answer Translation Module: Initial English-to-Galician translations of the medical answers were executed using SalamandraTA 7B Instruct.
  • Question–Answer Alignment & Synthetic Generation Module: To optimize question-to-answer semantic alignment, Gemma 3 27B Instruct was used to directly generate and align the enhanced questions in both English and Galician.
  • Expert Human Curation: To guarantee domain accuracy and linguistic integrity, human post-editing was conducted in compliance with ISO 18587:2017 standards and validated against official Galician medical references (Diccionario galego de termos médicos and Vocabulario de Medicina).

The repository includes a comprehensive documentation file along with three cross-aligned CSV datasets linked by a shared record identifier (original_id):

  • README.md: Contains exhaustive documentation regarding dataset usage, column descriptors and error taxonomy codes.
  • Primary Bilingual Parallel Corpus (GL_MedQuAD_bilingual_translations.csv): Includes raw English QA pairs, machine-translated Galician versions of the answers, ISO-curated Galician translations, synthetically generated questions in both English and Galician with curated Galician variants, and instruction-tuned response formats optimized for LLM fine-tuning and conversational healthcare applications.
  • Translation Quality Assessment Dataset (GL_MedQuAD_translation_evaluation.csv)
    Combines reference-less automated quality metrics (COMETKiwi, LaBSE) with expert human evaluation scores on a 0–5 scale, complemented by a fine-grained 12-category human post-editing error taxonomy.
  • Source-Text Linguistic Complexity Dataset (GL_MedQuAD_linguistic_complexity.csv)
    Captures the source-text linguistic complexity of the original English answers across three structured, complementary tiers, spanning human-perceived difficulty ratings provided by domain experts, composite complexity indices derived through Multi-Criteria Decision-Making (MCDM) frameworks—such as Analytic Hierarchy Process (AHP), CRITIC weighting, and Hybrid strategies—and fine-grained individual linguistic metrics.

Files

GL_MedQuAD_bilingual_translation.csv

Files (19.5 MB)

Name Size Download all
md5:f7279b73f95a364a5a9b0c112c843421
9.8 MB Preview Download
md5:1418252b2fcd9ffece7ae164b1e8d2ac
2.8 MB Preview Download
md5:f311d8b449bea0d36aad0695bd7299da
6.9 MB Preview Download
md5:536541bd830bc5abcdd35156d13d161e
13.0 kB Preview Download

Additional details

Funding

European Union
Xunta de Galicia
Axencia Galega de Innovacion
Universidade de Vigo
Consellería de Cultura, Educación, Formación Profesional e Universidades

Dates

Available
2026-07-31
Version v1