A Transformation-Based Benchmark for Evaluating the Robustness of LLMs in Generating OCL
Authors/Creators
Description
Replication Package for MODELS 2026 Research Track titled - A Transformation-Based Benchmark for Evaluating the Robustness of LLMs in Generating OCL
This repository provides a benchmarking pipeline for evaluating Large Language Models (LLMs) on Object Constraint Language (OCL) generation from:
- UML class diagrams (PlantUML format)
- Natural language specifications
It includes:
- Original UML/OCL datasets
- Natural language specification for each UML model
- 3 Systematic UML transformations - identifier renaming, attribute refication, and association refication
- Prompting framework for LLM-based OCL generation
The directory structure is as follows:├── dataset│ ├── UML│ │ ├── Airport│ │ │ ├── Airport.puml│ │ │ ├── Airport.use│ │ │ └── Airport.ocl│ │ └── EmploymentAgency│ ││ ├── Transformed_UML│ ├── specification.json│ └── transformed_specification.json │└── transformations│ ├── __init__.py│ ├── utils.py│ ├── uml_parser.py│ ├── rename_transformation.py│ ├── attribute_transformation.py│ ├── association_transformation.py│ └── runner.py├── evaluation│ ├── llm_runner.py│ ├── prompts.py│ ├── fine_tuning_LLM_for_OCL_script.ipynb│ └── config.py│└── run.py specification.json, and transformed_specification.json contain the natural language specification for each UML model
Fine-Tuning
fine_tuning_LLM_for_OCL_script.ipynb fine-tunes a causal LM on OCL generation using QLoRA (4-bit quantization + LoRA adapters).
- Dataset:
fpan/text-to-ocl-from-ecore(https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore), split 80/10/10 into train/val/test. - Requirements: HuggingFace account with an API token, and a GPU runtime (e.g. Google Colab). To run, open the notebook, replace
"HUGGINGFACE_API_TOKEN"with your token, and execute all cells.
Installation
https://doi.org/10.5281/zenodo.20454636pip install -r requirements.txtevaluation/config.py with the OpenRouter API key. Evaluation Procedure
Prerequisites
https://github.com/useocl/usepython run.py --mode {transformation} --task transformrename, attribute, association, rename_attribute, rename_association, attribute_association, fullpython run.py --mode rename --task transformpython run.py --mode {transformation} --task llmpython run.py --mode rename --task llmpython run.py --mode full --task bothdataset/UML/<ModelName>/<ModelName>.use.use file contains:open <test_instance>.soilpython run.py --mode {transformation} --task transformpython run.py --mode {transformation} --task llmAdditional Documentation
https://github.com/useocl/useReproducing the Experimental Results
python run.py --mode full --task transformpython run.py --mode full --task llmExample
python run.py --mode rename --task llmdataset/Transformed_renamed_UML/Airport/transformed_renamed_Airport.ocldataset/Transformed_renamed_UML/Airport/transformed_renamed_Airport.usedataset/Transformed_renamed_UML/Airport/Test_instances/*.soilSupplementary Material
Supplementary_Material.pdf provides real-world evidence grounding the three benchmark transformations
Files
Artifact.zip
Files
(813.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:13860346d8ec3aa8f8ac1cf66da6e412
|
813.8 kB | Preview Download |
Additional details
Dates
- Submitted
-
2026-03-27