Published April 8, 2021 | Version 1.0

Ecore Metamodels and EcoreBERT Pre-trained Language Model

Authors/Creators

  • 1. University of Montreal

Description

This dataset contains ecore metamodels from the MAR dataset transformed into tree representations. The original dataset can be found here: http://mar-search.org/experiments/models20/

The data contained in this repository were used to conduct the experiments in the paper: Recommending Metamodel Concepts during Modeling Activities with Pre-Trained Language Models. Link to the paper: https://arxiv.org/abs/2104.01642

The data are organized as follows:

  • model : our model trained on the tree representations of metamodels with RoBERTa architecture.
  • tokenizers : the byte-level BPE tokenizer we used to train our model.
  • train : the training data separated into a training and validation set.
  • test : the test data of all experiments conducted in the paper.

This data repository is linked with the following Github repository containing our code: https://github.com/martiwey/metamodel-concepts-bert

Files

train_vocab.txt

Files (1.2 GB)

Name Size
md5:780a11acf6b088063f72ff6190fe621f
1.2 GB Download
md5:70f3fbbbd0d5b92712200abc49bc97b1
154.8 kB Download
md5:ce08895d38aea64b5517fcb218af5e4d
272.8 kB Download
md5:cf31149908be1fda01d199c5fccb5a52
826.9 kB Download
md5:8ec8815c1f653ac285f73e662730d62b
861.5 kB Preview Download