Published August 21, 2026 | Version 1.1.0

DISSILEX: A lexico-semantic network and valency lexicon for medieval Latin

Contributors

Data collector:

Data curator:

  • 1. ROR icon University of Antwerp
  • 2. ROR icon Masaryk University
  • 3. ROR icon Czech Academy of Sciences

Description

DISSILEX is a controlled vocabulary in the form of a manually built lexico-semantic network of medieval Latin verbs and verbal expressions, featuring a detailed valency lexicon and connections to a large set of Latin and modern-English concepts, presented here as a single SQLite file (dissilex.db) that is readable by the sqlite3 command-line tool, any SQLite browser, or Python's built-in sqlite3 module.

Rooted in the domain of inquisitorial records, DISSILEX covers general as well as more subject-specific meanings, with both standard (synonym, hypernym, etc.) and less canonical relations. DISSILEX is a product of Computer-Assisted Semantic Text Modelling (CASTEMO; Zbíral et al. 2026 - see README.md for full references), an approach to modeling statements as a four-slot structure of subject(s), predicate(s) and two objects, creating a thickly connected network of data points. Coverage is richest for human-interaction verbs (testimony, accusation, belief, religious practice) and the legal vocabulary of heresy trials, making it a machine-operable resource for modelling Latin textual data on dissent, resistance, repression, and resilience.

We distinguish two entry types: Actions (verbs and verbal expressions, each associated with a three-slot valency frame specifying entity type, morphosyntactic, and semantic valencies) and Concepts (single- and multi-word expressions for other parts of speech). The network is connected through a set of 11 relation types, including superclass (hypernym) membership, synonymy, antonymy, verb-to-noun mappings, and valency-specific relations. Each relation connects two entities, and can be unidirectional or bidirectional.

As part of an ongoing effort to position DISSILEX within the Linguistic Linked Open Data (LLOD) cloud, many entries contain IDs to external sources stored in the database, specifically to the LiLa (Linking Latin) Lemma Bank and Princeton WordNet (PWN) 3.0 and 3.1 synsets. We applied the Collaborative Interlingual Index (CILI) to map between the two versions of the PWN for entries where only one of the IDs has been added. We also indicate cases where no equivalent for a DISSILEX lemma exists ("NA").

Via the LiLa SPARQL endpoint, it is possible to use the linked LiLa lemmas, which feature as the central unit of linking sources in the Latin LLOD cloud, to retrieve data from several resources including dictionaries, corpora, treebanks, and various NLP tools. We have made use of this opportunity to enrich the database file with lemmas from the LiLa Lemma Bank, while also supplying LatinCy-generated lemmas for most Actions (model: la_core_web_lg).

This release contains:

  • dissilex.db: SQLite database, which can be readily queried
  • dissilex_schema.md / dissilex_schema.pdf: schema documentation
  • README.md: full dataset description, statistics, and SQL examples
  • ATTRIBUTION.md: license and attribution notices.
  • LICENSE-DATA: Full CC BY-SA 4.0 license text.

Funding, attribution and licence

DISSILEX is developed by the Dissident Networks research group (DISSINET) at Masaryk University and has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme, grant agreement No. 101000442, project “Networks of Dissent: Computational Modelling of Dissident and Inquisitorial Cultures in Medieval Europe”, and from the European Regional Development Fund, grant agreement No. CZ.02.01.01/00/22_008/0004595, project “Beyond Security: Role of Conflict in Resilience-Building”. 

DISSILEX is released under CC BY-SA 4.0. It incorporates data from external resources:

  • LiLa Lemma Bank: CIRCSE, Università Cattolica del Sacro Cuore (Milan). Licensed CC BY-SA 4.0. The database redistributes a subset of LiLa lemma forms (subset-selected and format-converted, not otherwise modified); the ShareAlike clause is honoured by this release's CC BY-SA 4.0 licence. URL: https://lila-erc.eu/
  • Princeton WordNet 3.0 / 3.1: We distribute offset IDs for both versions, and gloss text from WordNet 3.0 only. (An entry with a 3.1 identifier carries the corresponding 3.0 gloss, mapped via CILI.) WordNet 3.0 Copyright 2006 by Princeton University. All rights reserved. WordNet License. https://wordnet.princeton.edu/
  • LatinCy/spaCy: We redistribute output from the LatinCy model. The `spacy_lemma` field contains lemmas generated by the LatinCy spaCy pipeline `la_core_web_lg` (Patrick J. Burns). The model is MIT-licensed (<https://huggingface.co/latincy/la_core_web_lg/blob/main/README.md>); the spaCy library is MIT-licensed (<https://github.com/explosion/spaCy/blob/master/LICENSE>).

Full notices, including the Collaborative Inter-Lingual Index (CILI) and Latin WordNet, are in ATTRIBUTION.md.

Version 1.1.0 is mostly a quality improvement of existing entries, but also adds and removes entries and links to external resources. The schema is unchanged, so queries written against version 1.0.0 continue to work.

Files

ATTRIBUTION.md

Files (3.8 MB)

Name Size Download all
md5:eeb1789585ff032240dee3442ce55d6f
7.9 kB Preview Download
md5:70871736bcd8ebf387ec8568090c7064
3.6 MB Download
md5:ec5a1aeba494165f983fc5ccde46ec7d
7.1 kB Preview Download
md5:8f1f164cc102288fbbeb3386e525e09b
63.3 kB Preview Download
md5:22449197f5884b3a25aac08965bd83a3
20.1 kB Download
md5:9dceb572e8c35daa551a71333856a0b3
19.1 kB Preview Download

Additional details

Related works

Is part of
Dataset: 10.5281/zenodo.17936848 (DOI)

Funding

European Commission
DISSINET - Networks of Dissent: Computational Modelling of Dissident and Inquisitorial Cultures in Medieval Europe 101000442
European Commission
Beyond Security: Role of Conflict in Resilience-Building CZ.02.01.01/00/22_008/0004595

Dates

Available
2026-08-21
1.1.0 public release

Software

Repository URL
https://github.com/DISSINET/dissilex-app
Programming language
Python , SQL
Development Status
Active

References

  • Baker, C. F., Fillmore, C. J., and Lowe, J. B. (1998): "The Berkeley FrameNet project". In: Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics, Volume 1. Montreal: Association for Computational Linguistics, pp. 86–90. DOI: 10.3115/980845.980860
  • Bond, F, Vossen, P, McCrae, J. P., and Fellbaum, C. (2016): "CILI: the Collaborative Interlingual Index". In: Fellbaum, C., Vossen, P., Mititelu, V. B., and Forascu, C. (eds.): Proceedings of the 8th Global WordNet Conference (GWC). Bucharest: Global Wordnet Association, pp. 50–57. DOI: 10.18653/v1/2016.gwc-1.9
  • McCrae, J. P., Bosque-Gil, J., Gracia, J., Buitelaar, P., and Cimiano, P. (2017): "The Ontolex-Lemon model: development and applications". In: Proceedings of eLex 2017 conference. Leiden, 587–597.
  • Chiarcos, C., McCrae, J. P., Cimiano, P., and Fellbaum, C. (2013): "Towards Open Data for Linguistics: Linguistic Linked Data". In: Oltramari, A., Vossen, P, Qin, L., and Hovy, E. (eds.): New Trends of Research in Ontologies and Lexical Resources. Ideas, Projects, Systems. Berlin / Heidelberg: Springer, pp. 7–25. DOI: 10.1007/978-3-642-31782-8_2
  • Cimiano, P., Chiarcos, C., McCrae, J. P., and Jorge, G. (2020): "Linguistic Linked Open Data Cloud". In: Cimiano, P., Chiarcos, C., McCrae, J. P., and Gracia, J. (eds.): Linguistic Linked Data: Representation, Generation and Applications. Cham: Springer International Publishing, 29–41. DOI: 10.1007/978-3-030-30225-2
  • Corcoglioniti, F., Rospocher, M., Aprosio, A. P., and Tonelli, S. (2016): "PreMOn: a Lemon Extension for Exposing Predicate Models as Linked Data". In: Calzolari, N. et al. (eds.): Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16). Portorož: European Language Resources Association (ELRA), 877–884. https://aclanthology.org/L16-1141/
  • Mambrini, F. and Passarotti, M. C. (2023): "The LiLa Lemma Bank: A Knowledge Base of Latin Canonical Forms". In: Journal of Open Humanities Data 9, 1. DOI: 10.5334/johd.145
  • Miller, G. A. (1995): "WordNet: a lexical database for English". In: Communications of the ACM 38, 11: 39–41. DOI: 10.1145/219717.219748
  • Minozzi, S. (2017): "Latin WordNet, una rete di conoscenza semantica per il latino e alcune ipotesi di utilizzo nel campo dell'information retrieval", in: Mastandrea, P. (ed.): Strumenti digitali e collaborativi per le Scienze dell'Antichità. Venezia: Edizioni Ca' Foscari (Antichistica 14). DOI: 10.14277/6969-182-9/ANT-14-10
  • Navigli, R. and Ponzetto, S. P. (2012): "BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network". In: Artificial Intelligence 193: 217–250. DOI: 10.1016/j.artint.2012.07.001
  • Passarotti, M. C., Mambrini, F., Franzini, G., Cecchini, F. M., Litta, E. Moretti, G., Ruffolo, P., and Sprugnoli, R. (2020): "Interlinking through Lemmas. The Lexical Collection of the LiLa Knowledge Base of Linguistic Resources for Latin". In: Studi e Saggi Linguistici 58, 1: 177–212. DOI: 10.4454/ssl.v58i1.277.
  • Passarotti, M. C., Saavedra, B. G., and Onambele, C. (2016): "Latin Vallex. A Treebank-based Semantic Valency Lexicon for Latin". In: Calzolari, N., Choukri, K., Declerck, T., Goggi, S., Grobelnik, M., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., Odijk, J., and Piperidis, S. (eds.): Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16). Portorož: European Language Resources Association (ELRA), 2599–2606. https://aclanthology.org/L16-1414/
  • Schuler, K. K. (2005): VerbNet: A broad-coverage, comprehensive verb lexicon. Ph.D. thesis, University of Pennsylvania. https://www.proquest.com/openview/7ca4b1b9093522a7d8089ff2e987e74e/
  • Van Valin Jr., R. D. and Foley, W. A. (1980): "Role and Reference Grammar". In: Moravcsik, E. A. and Wirth, J. R. (eds.): Syntax and Semantics 13: Current Approaches to Syntax. New York: Academic Press, 329–352. DOI: 10.1163/9789004373105_014
  • Žabokrtský, Z. (2005): Valency lexicon of Czech verbs. Ph.D. thesis, Charles University, Prague.
  • Zbíral, D., Shaw, R. L. J., Hanák, P., Hampejs, T., and Mertel, A. (2025): From texts to structured data: Building knowledge graphs through Computer-Assisted Semantic Text Modelling (CASTEMO). Brno: Masaryk University. URL: https://docs.religionistika.phil.muni.cz/books/from-texts-to-structured-data-building-knowledge-graphs-through-computer-assisted-semantic-text-modelling-castemo
  • Fellbaum, C. (1998) WordNet: An electronic lexical database. MIT press.
  • Princeton University. (2010) What is WordNet? URL: https://wordnet.princeton.edu/. Accessed: 2 June 2026.
  • Burns, P. J. (2023). LatinCy: Synthetic trained pipelines for Latin NLP. arXiv. https://doi.org/10.48550/arXiv.2305.04365