There is a newer version of the record available.

Published February 26, 2022 | Version v2

A Systematic Evaluation of Large Language Models of Code

  • 1. Carnegie Mellon University

Description

These are datasets for the paper:

"A Systematic Evaluation of Large Language Models of Code"

https://arxiv.org/pdf/2202.13169.pdf

The code is available at: https://github.com/VHellendoorn/Code-LMs

 

The file "unseen_test_sets.tar.gz" contains test sets of ~100 files in each of 12 programming languages.

These files are not included in The Pile, and thus models such as GPT-Neo, GPT-J, GPT-NeoX were not trained on them.

In the paper, we use these test sets to compare a variety of language models of code including OpenAI's Codex, GPT-J, GPT-Neo, GPT-NeoX-20B, and CodeParrot and our PolyCoder model.

 

The file "index.zip" includes an index of the training set file paths and commit SHAs.

Notes

https://arxiv.org/abs/2202.13169

Files

index.zip

Files (1.3 GB)

Name Size
md5:eaf63ec6e69478bde42b1b28a46a1881
1.3 GB Preview Download
md5:8d8b890d2cfa7d26530b028f1a6fa47e
904.1 kB Download