Merging DanNet with Princeton Wordnet
Creators
- 1. University of Copenhagen
- 2. The Danish Society for Language and Literature
Description
In this paper we describe the merge of the Danish wordnet, DanNet, with Princeton Wordnet applying a two-step approach. We first link from the English Princeton core to Danish (5,000 base concepts) and then proceed to link-ing the rest of the Danish vocabulary to English, thus going from Danish to English. Since the Danish wordnet is built bottom-up from Danish lexica and corpora, all taxonomies are monolingually based and thus not necessarily directly compatible with the coverage and structure of the Princeton WordNet. This fact proves to pose some challenges to the linking procedure since a considerable number of the links cannot be realised via the preferred cross-language synonym link which implies a more or less precise correlation between the two concepts. Instead, a subpart of the links are realised through near synonym or hyponymy links to compensate for the fact that no precise translation can be found in the target resource. The tool WordnetLoom is currently used for manual linking but procedures for a more automatic procedure in future is discussed. We conclude that the two resources actually differ from each other quite more than expected, both vocabulary- and structure-wise.
Files
LinkingDanNettoPrinceton_final_26-06.pdf
Files
(700.9 kB)
Name | Size | Download all |
---|---|---|
md5:377edfdc9f7d3885b41fc0c19e7b0b13
|
700.9 kB | Preview Download |