Published May 31, 2024 | Version v1
Dataset Open

Analyzing the Dependability of Large Language Models for Code Clone Generation.

Contributors

Contact person:

Description

data.zip: 

This dataset includes a collection of ten LeetCode programming problems used in the study "Analyzing the Dependability of Large Language Models for Code Clone Generation". At the top level, you will find a CSV file containing all the initial LeetCode data. Each subdirectory at this level represents a specific LeetCode problem. Within these subdirectories, you will find the original solutions, their behaviors, the input corpus, as well as folders dedicated to various temperatures, models, and code cloning tasks. Additionally, within the "repeated" folder, you will find the original LLM-generated snippets, the preprocessed snippets with the snippet behavior, and the results.

characterizing_code_clones_project.zip: 

This zipped directory encompasses the core scripts and results used in the "Characterizing Code Clones of LLMs" research. The common folder was used to run the whole pipeline, the various parts of the pipeline are each in a folder as well as the various data analysis scripts! 

 

Files

data.zip

Files (58.7 MB)

Name Size Download all
md5:639c1b5f594facc2492ea5dfd81e2c96
55.4 MB Preview Download
md5:e95e6c3dcf0a9f690b1c5a4a4e102372
3.3 MB Preview Download