Dataset Open Access

Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation

Panchenko, Alexander; Ruppert, Eugen; Faralli, Stefano; Ponzetto, Simone Paolo; Biemann, Chris

This dataset contains the models for interpretable Word Sense Disambiguation (WSD) that were employed in Panchenko et al. (2017; the paper can be accessed at

The files were computed on a 2015 dump from the English Wikipedia. Their contents:

  • Induced Sense Inventories: wp_stanford_sense_inventories.tar.gz
    This file contains 3 inventories (coarse, medium fine)
  • Language Model (3-gram):
    This file contains all n-grams up to n=3 and can be loaded into an index
  • Weighted Dependency Features: wp_stanford_lemma_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000.gz
    This file contains weighted word--context-feature combinations and includes their count and an LMI significance score
  • Distributional Thesaurus (DT) of Dependency Features: wp_stanford_lemma_BIM_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000_simsortlimit200_feature expansion.gz
    This file contains a DT of context features. The context feature similarities can be used for context expansion

For further information, consult the paper and the companion page:

Panchenko A., Ruppert E., Faralli S., Ponzetto S. P., and Biemann C. (2017): Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL'2017). Valencia, Spain. Association for Computational Linguistics.



Files (10.6 GB)
Name Size
9.5 GB Download
wp_stanford_lemma_BIM_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000_simsortlimit200_feature expansion.gz
344.1 MB Download
404.5 MB Download
342.8 MB Download
All versions This version
Views 7070
Downloads 1818
Data volume 43.2 GB43.2 GB
Unique views 7070
Unique downloads 1010


Cite as