Published February 18, 2018 | Version 1.0

Deep Reference Mining from Scholarly Literature in the Arts and Humanities - Pre-trained word embeddings

Authors/Creators

  • 1. The Alan Turing Institute

Contributors

Researcher:

  • 1. Matteo

Description

Pre-trained word vectors of dimensionality 100 and 300 for the publication: Deep Reference Mining from Scholarly Literature in the Arts and Humanities, submitted to Frontiers in Digital Humanities.

The corpus of scholarly publications from which these vectors were trained is under copyright, therefore we publish these vectors for reproducibility. Please refer to the publication's repository for further details: https://github.com/dhlab-epfl/LinkedBooksDeepReferenceParsing.

These vectors were trained using Gensim 3.1.0. The corpus was preprocessed as follows:

  1. word tokenization with NLTK word_punct tokenizer.
  2. digits were converted into the $NUM$ token
  3. words less frequent than 5 times, for every document, were converted to the $UNK$ token
  4. vectors were trained using the function: Word2Vec(window=5, min_count=5, sg=1)

Files

Files (1.1 GB)

Name Size
md5:f091934635c5772b41866fd2eb26668a
1.1 GB Download

Additional details

Funding

Swiss National Science Foundation
Linked Books: Reconstructing the history of the history of Venice 205121_159961