Unsupervised keyword extraction using the GoW model and centrality scores

doi:10.1007/978-3-319-70284-1_26

Published February 13, 2018 | Version v1

Conference paper Open

Unsupervised keyword extraction using the GoW model and centrality scores

1. Aristotle University of Thessaloniki, Greece
2. CERTH-ITI, Thessaloniki, Greece

Nowadays, a large amount of text documents are produced on a daily basis, so we need efficient and effective access to their content. News articles, blogs and technical reports are often lengthy, so the reader needs a quick overview of the underlying content. To that end we present graph-based models for keyword extraction, in order to compare the Bag of Words model with the Graph of Words model in the keyword extraction problem. We compare their performance in two publicly available datasets using the evaluation measures Precision@10, mean Average Precision and Jaccard coefficient. The methods we have selected for comparison are grouped into two main categories. On the one hand, centrality measures on the formulated Graph-of-Words (GoW) are able to rank all words in a document from the most central to the less central, according to their score in the GoW representation. On the other hand, community detection algorithms on the GoW provide the largest community that contains the key nodes (words) in the GoW. We selected these
methods as the most prominent methods to identify central nodes in a GoW model. We conclude that term-frequency scores (BoW model) are useful only in the case of less structured text, while in more structured text documents, the order of words plays a key role and graph-based
models are superior to the term-frequency scores per document.

Files

unsupervised-keyword-extraction-camera-ready.pdf

Files (624.1 kB)

Name	Size	Download all
unsupervised-keyword-extraction-camera-ready.pdf md5:c9a9d3131b2d81529061700b5e6aae06	624.1 kB	Preview Download

Additional details

KRISTINA – Knowledge-Based Information Agent with Social Competence and Human Interaction Capabilities 645012: European Commission
TENSOR – Retrieval and Analysis of Heterogeneous Online Content for Terrorist Activity Recognition 700024: European Commission

	All versions	This version
Views	160	160
Downloads	91	91
Data volume	59.3 MB	59.3 MB

Unsupervised keyword extraction using the GoW model and centrality scores

Creators

Description

Files

unsupervised-keyword-extraction-camera-ready.pdf

Files (624.1 kB)

Additional details

Funding