Published November 29, 2022 | Version 1

Bangla Information Retrieval Test Collection | Revisiting Anwesha

  • 1. Indian Institute of Technology Madras
  • 2. Centre for Development of Advanced Computing, Kolkata
  • 3. indian Institute of Technology Madras

Description

There are several IR test collections available in English (e.g. http://ir.dcs.gla.ac.uk/resources/test_collections/). Unfortunately, no Gold standard dataset existed for Bangla IR until recently (https://zenodo.org/record/6583149). Our work expands the existing Gold standard dataset by creating 100 query document relevance pairs across a new test collection of 1000 documents. The corpus contains news articles from Ebela, Zee News and Anandabazar Patrika, Vikaspedia and various Bangla travel blogs. The definition of the complexity level of a query is described below:

Complexity Level 1: The query contains exact words, phrases or sentence from the document.

Complexity Level 2: The query is not present as it is in the document. There is a slight deviation.

Complexity Level 3: The query is a generalised phrase capturing the overall story or the document’s theme.

Complexity Level 4: It is a general query not related to any specific document.

Files

BSE_qrels.json

Files (12.8 MB)

Name Size Download all
md5:d3054dade591424a81d0c7fe8e3be568
19.7 kB Preview Download
md5:0f43b47c06b82040bf4c1a0e6585f622
10.7 MB Preview Download
md5:2d3ce91c3034a09e3f5e86fa67dd3b5a
2.1 MB Preview Download
md5:78cde749bb6bd39482a9965d4b75ecc5
36.5 kB Download