Published May 26, 2020 | Version 1.0

Experimental Data for "What Makes a Top-Performing Precision Medicine Search Engine? Tracing Main System Features in a Systematic Way" at SIGIR2020

  • 1. Jena University Language and Information Engineering (JULIE) Lab, Friedrich-Schiller-University Jena
  • 2. Institute for Medical Informatics, Statistics and Documentation, Medical University of Graz

Description

This deposit contains data used for the experiments reported in the paper "What Makes a Top-Performing Precision Medicine Search Engine? Tracing Main System Features in a Systematic Way", most notably the ElasticSearch 5.4 indices used for the reported experiments.

To load the indices into an ElasticSearch cluster of your own, use the restore function described in the ElasticSearch documentation.

The names of the index snapshots contained here are

  • ct1718 for the indexed ClinicalTrials data used in the TREC-PM challenges in 2017 and 2018.
  • ct19 for the indexed ClinicalTrials data used in the TREC-PM challenge in 2019.
  • ba1718 for the indexed PubMed data used in the TREC-PM challenges in 2017 and 2018.
  • ba19 for the indexed PubMed data used in the TREC-PM challenge in 2019.

The other file contains the original output that SMAC wrote to disc during the parameter optimization process. There are directories for the biomedical abstracts (BA) and clinical trials (ct) and for each respective 10 fold cross validation split. Those file contain the exact parameter configurations and their evalation score (the infNDCG metric was used) in live-runXX.json files.

The code to these files is located in this Zenodo deposit.

Notes

This work was supported by the BMBF within the SMITH project under grant 01ZZ1803G.

Files

elasticsearch-indices.zip

Files (106.0 GB)

Name Size
md5:c483cd25d82462b85719f9c6c34fa63b
105.8 GB Preview Download
md5:522ffbeef138b41eaedce40a2bf9c94d
194.7 MB Preview Download