Published September 1, 2023 | Version v1

Document retrieval using term frequency inverse sentence frequency weighting scheme

  • 1. College of Health and Medical Technology, Middle Technical University, Baghdad, Iraq
  • 2. College of Science, University of Baghdad, Baghdad, Iraq

Description

The need for an efficient method to find the furthermost appropriate document corresponding to a particular search query has become crucial due to the exponential development in the number of papers that are now readily available to us on the web. The vector space model (VSM) a perfect model used in “information retrieval”, represents these words as a vector in space and gives them weights via a popular weighting method known as term frequency inverse document frequency (TF-IDF). In this research, work has been proposed to retrieve the most relevant document focused on representing documents and queries as vectors comprising average term term frequency inverse sentence frequency (TF-ISF) weights instead of representing them as vectors of term TF-IDF weight and two basic and effective similarity measures: Cosine and Jaccard were used. Using the MS MARCO dataset, this article analyzes and assesses the retrieval effectiveness of the TF-ISF weighting scheme. The result shows that the TF-ISF model with the Cosine similarity measure retrieves more relevant documents. The model was evaluated against the conventional TF-ISF technique and shows that it performs significantly better on MS MARCO data (Microsoft-curated data of Bing queries).

Files

30532-67223-1-PB.pdf

Files (437.9 kB)

Name Size Download all
md5:cd4fe5db96f4dddae5029eb9c2086eef
437.9 kB Preview Download