Published May 31, 2024 | Version v1
Dataset Open

Analyzing the Dependability of Large Language Models for Code Clone Generation

Authors/Creators

Description

data.zip: 

This dataset includes a collection of ten LeetCode programming problems used in the study "Analyzing the Dependability of Large Language Models for Code Clone Generation". At the top level, you will find a CSV file containing all the initial LeetCode data. Each subdirectory at this level represents a specific LeetCode problem. Within these subdirectories, you will find the original solutions, their behaviors, the input corpus, as well as folders dedicated to various temperatures, models, and code cloning tasks. Additionally, within the "repeated" folder, you will find the original LLM-generated snippets, the preprocessed snippets with the snippet behavior, and the results.

characterizing_code_clones_project.zip: 

This zipped directory encompasses the core scripts and results used in the "Characterizing Code Clones of LLMs" research. The common folder was used to run the whole pipeline, the various parts of the pipeline are each in a folder as well as the various data analysis scripts! 

Files

data.zip

Files (59.1 MB)

Name Size Download all
md5:2a36686a6a3b8d0d4fd0c455ccab97e7
55.8 MB Preview Download
md5:e95e6c3dcf0a9f690b1c5a4a4e102372
3.3 MB Preview Download