Published October 18, 2025 | Version v1

Question Benchmark Dataset for Valencia Tourism information RAG

Description

Dataset Description

This repository presents a domain-specific question-answering (QA) dataset designed for Retrieval-Augmented Generation (RAG) applications in the tourism domain, with a focus on Valencia, Spain. The dataset comprises 994 question-answer pairs grounded in a curated corpus of 84 Spanish-language documents pertaining to Valencia's tourism information.

Corpus Composition

The textual corpus has been systematically compiled from authoritative sources, including:

  • Tourism-oriented blog articles
  • Wikipedia entries related to Valencia
  • Official tourism guides published by VisitValencia

All source materials are provided in Spanish and have undergone standardized preprocessing procedures, including conversion to plain text format and coreference resolution to normalize entity mentions throughout the corpus.

Dataset Structure

The repository contains:

  1. documents.zip: Complete collection of 84 curated source documents
  2. QA dataset (dataset.csv): 994 question-answer pairs, each annotated with:
    • Context window (relevant text excerpt containing the answer)
    • Source document identifier

Generation Methodology

The question-answer pairs were generated programmatically using Google's Gemini 2.5 Flash model, ensuring each QA pair is contextually grounded in the source documentation. This approach guarantees factual consistency between generated questions, answers, and supporting textual evidence.

Intended Applications

This dataset is specifically designed for the development, training, and evaluation of domain-specific RAG systems targeting tourism information retrieval and question-answering tasks in Spanish-language contexts.

Files

dataset.csv

Files (906.5 kB)

Name Size Download all
md5:93ec2f4840857e56b141a38b96ab32f0
496.3 kB Preview Download
md5:f9556bdd5b3cdc32f145c11457052303
410.1 kB Preview Download