Published July 31, 2026
| Version v1
Dataset
Open
GL-MedQuAD: A Curated English–Galician Medical Question Answering Dataset
Authors/Creators
Description
This data descriptor introduces GL-MedQuAD, the first domain-specific biomedical Question Answering (QA) benchmark for the Galician language. Comprising 2,100 parallel records derived from the MedQuAD corpus—specifically from GHR (N=1,467) and NIHSeniorHealth (N=633)—the dataset addresses the critical scarcity of specialized clinical NLP resources in low-resource linguistic scenarios.
The dataset was constructed through a hybrid, multi-stage generation-and-curation pipeline:
- Answer Translation Module: Initial English-to-Galician translations of the medical answers were executed using SalamandraTA 7B Instruct.
- Question–Answer Alignment & Synthetic Generation Module: To optimize question-to-answer semantic alignment, Gemma 3 27B Instruct was used to directly generate and align the enhanced questions in both English and Galician.
- Expert Human Curation: To guarantee domain accuracy and linguistic integrity, human post-editing was conducted in compliance with ISO 18587:2017 standards and validated against official Galician medical references (Diccionario galego de termos médicos and Vocabulario de Medicina).
The repository includes a comprehensive documentation file along with three cross-aligned CSV datasets linked by a shared record identifier (original_id):
- README.md: Contains exhaustive documentation regarding dataset usage, column descriptors and error taxonomy codes.
- Primary Bilingual Parallel Corpus (GL_MedQuAD_bilingual_translations.csv): Includes raw English QA pairs, machine-translated Galician versions of the answers, ISO-curated Galician translations, synthetically generated questions in both English and Galician with curated Galician variants, and instruction-tuned response formats optimized for LLM fine-tuning and conversational healthcare applications.
- Translation Quality Assessment Dataset (GL_MedQuAD_translation_evaluation.csv)
Combines reference-less automated quality metrics (COMETKiwi, LaBSE) with expert human evaluation scores on a 0–5 scale, complemented by a fine-grained 12-category human post-editing error taxonomy. - Source-Text Linguistic Complexity Dataset (GL_MedQuAD_linguistic_complexity.csv)
Captures the source-text linguistic complexity of the original English answers across three structured, complementary tiers, spanning human-perceived difficulty ratings provided by domain experts, composite complexity indices derived through Multi-Criteria Decision-Making (MCDM) frameworks—such as Analytic Hierarchy Process (AHP), CRITIC weighting, and Hybrid strategies—and fine-grained individual linguistic metrics.
Files
GL_MedQuAD_bilingual_translation.csv
Additional details
Funding
- European Union
- Xunta de Galicia
- Axencia Galega de Innovacion
- Universidade de Vigo
- Consellería de Cultura, Educación, Formación Profesional e Universidades
Dates
- Available
-
2026-07-31Version v1