Published January 27, 2020
| Version v1
Dataset
Open
Raw C Code Corpus
Authors/Creators
- 1. The University of Edinburgh
- 2. Free University of Bozen-Bolzano
- 3. Google Research
Description
A raw code corpus for the C programming language i.e., includes only the C source files of each repository without any preprocessing.
The corpus was used to generate the C training, validation, testing, and BPE encoding sets for the experiments performed in the paper: Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code.
Files
Additional details
Related works
- Is source of
- Dataset: 10.5281/zenodo.3628638 (DOI)
- Other: 10.5281/zenodo.3628628 (DOI)
Funding
- UK Research and Innovation
- EPSRC Centre for Doctoral Training in Data Science EP/L016427/1