IR Lab Cologne/Jena/Kassel Winter Term 2024/2025
- Zayene, Ahmed
- Kumarasamy, Ajeeth
- Otsepaev, Alex
- Meuche, André Philipp
- Kiafard, Ashkan
- Stein, Benno
- Bhagyasha, Patil
- Neumann, Dennis
- Kharsah, Einaas
- Naeem, Habiba
- Scells, Harrisen
- Kashin, Igor
- Qayas, Ilaf
- Jan Heinrich, Merker
- Hart, Johannes
- Kremer, Johannes
- Kacem, Khaireddine Hadj
- Kanojia, Kishanlal
- Böttjer, Kristan
- Hoffmann, Lea-Sophie
- Pickartz, Lena
- Obitz, Lucas
- Čulina, Luka
- Günther, Lukas
- Fröbe, Maik
- Hagen, Matthias
- Potthast, Martin
- Tennert, Matti
- Rissiek, Nele
- Mayer, Nico
- Alshehabi, Omran
- Stahl, Patrick
- Schaer, Philipp
- Werner, Quirin
- Hofmann, Richard
- Ahmed, Shehroz
- Löber, Shirina
- Wallraven, Stephan
- Shah, Syed Ghazanfar Ali
- Billerbeck, Till
- Köhne, Tim
- Hagen, Tim
- Tammaa, Ward
Description
The Datasets for the Information Retrieval Courses in Cologne/Jena/Kassel in Winter Term 2024/2025
This repository contains resources coupled to ir_datasets and TIREx for IR courses that focus their hands-on labs on shared tasks. During the IR exercises in winter term 2023/2024, we collaboratively developed and evaluated IR systems in a shared task style setup, covering corpus creation, system development, and statistical analysis. The resulting artifacts, i.e., the documents, topics, runs, relevance judgments can be browsed at https://tira.io/task-overview/ir-lab-wise-2024. This zenodo artifact contains all of the underlying datasets used and produced during the course together with instructions on how to easily access the data using ir_datasets.
The artifact in this dataset include the following files:
- subsampled-ms-marco-deep-learning-20241201-training-inputs.zip containing the training inputs, i.e., containing the document corpus and the topics.
- subsampled-ms-marco-deep-learning-20241201-training-truths.zip containing the training truth to evaluate and tune systems, i.e., the topics and relevance judgments.
Accessing the Data with ir_datasets
We provide wrapper code to easily access the resources with ir_datasets:
# this loads a patched version of ir_datasets that can load resources from TIRA
from tira.third_party_integrations import ir_datasets
training_dataset = ir_datasets.load('ir-lab-wise-2024/subsampled-ms-marco-deep-learning-20241201-training')
Similarly, the same is possible with the ir_datasets integration to PyTerrier:
from tira.third_party_integrations import ensure_pyterrier_is_loaded
import pyterrier as pt
# this patches ir_datasets and loads PyTerrier so that it can load resources from TIRA and can run in the TIRA sandbox
ensure_pyterrier_is_loaded()
training_dataset = pt.datasets.get_dataset('irds:ir-lab-wise-2024/subsampled-ms-marco-deep-learning-20241201-training')
Files
subsampled-ms-marco-ir-lab-20250105-test-inputs.zip
Files
(204.8 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:a1e21a40cc77a358b1cc17f9f0fc17ec
|
10.0 MB | Preview Download |
|
md5:ba9ed37048e4548892c2c15b3ecfb4ae
|
63.2 kB | Preview Download |
|
md5:9a614483bdf2a5c0d5f8531af1f2b823
|
111.2 MB | Preview Download |
|
md5:0e824ed2fe7e6c99b5d3acacc113027f
|
51.8 kB | Preview Download |
|
md5:5b3e5a720ec789e3fa3c9be3e1ebd277
|
83.6 MB | Preview Download |
|
md5:871ed918b9e2c14cc6f1295ff843d7d1
|
7.8 kB | Preview Download |