Question Benchmark Dataset for Valencia Tourism information RAG
Authors/Creators
Description
Dataset Description
This repository presents a domain-specific question-answering (QA) dataset designed for Retrieval-Augmented Generation (RAG) applications in the tourism domain, with a focus on Valencia, Spain. The dataset comprises 994 question-answer pairs grounded in a curated corpus of 84 Spanish-language documents pertaining to Valencia's tourism information.
Corpus Composition
The textual corpus has been systematically compiled from authoritative sources, including:
- Tourism-oriented blog articles
- Wikipedia entries related to Valencia
- Official tourism guides published by VisitValencia
All source materials are provided in Spanish and have undergone standardized preprocessing procedures, including conversion to plain text format and coreference resolution to normalize entity mentions throughout the corpus.
Dataset Structure
The repository contains:
- documents.zip: Complete collection of 84 curated source documents
- QA dataset (dataset.csv): 994 question-answer pairs, each annotated with:
- Context window (relevant text excerpt containing the answer)
- Source document identifier
Generation Methodology
The question-answer pairs were generated programmatically using Google's Gemini 2.5 Flash model, ensuring each QA pair is contextually grounded in the source documentation. This approach guarantees factual consistency between generated questions, answers, and supporting textual evidence.
Intended Applications
This dataset is specifically designed for the development, training, and evaluation of domain-specific RAG systems targeting tourism information retrieval and question-answering tasks in Spanish-language contexts.
Files
dataset.csv
Files
(906.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:93ec2f4840857e56b141a38b96ab32f0
|
496.3 kB | Preview Download |
|
md5:f9556bdd5b3cdc32f145c11457052303
|
410.1 kB | Preview Download |