There is a newer version of the record available.

Published November 10, 2021 | Version v2

Multilingual training for Software Engineering

Authors/Creators

Description

This is the replication package for "Multilingual training for Software Engineering".

Code Summarization

CodeBERT: for monolingual fine-tuning and inference, please clone the "CodeXGLUE" repo and follow the instruction in the given link

https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text

For multilingual fine-tuning and inference, copy the combine.zip folder to dataset folder and unzip it. Run the following commands:

cd code
lang=combine #programming language
lr=5e-5
batch_size=32
beam_size=10
source_length=256
target_length=128
data_dir=../dataset
output_dir=model/$lang
train_file=$data_dir/$lang/train.jsonl
dev_file=$data_dir/$lang/valid.jsonl
epochs=10
pretrained_model=microsoft/codebert-base #Roberta: roberta-base

python run.py --do_train --do_eval --model_type roberta --model_name_or_path $pretrained_model --train_filename $train_file --dev_filename $dev_file --output_dir $output_dir --max_source_length $source_length --max_target_length $target_length --beam_size $beam_size --train_batch_size $batch_size --eval_batch_size $batch_size --learning_rate $lr --num_train_epochs $epochs

batch_size=64
dev_file=$data_dir/$lang/valid.jsonl
test_file=$data_dir/$lang/test.jsonl
test_model=$output_dir/checkpoint-best-bleu/pytorch_model.bin #checkpoint for test

python run.py --do_test --model_type roberta --model_name_or_path microsoft/codebert-base --load_model_path $test_model --dev_filename $dev_file --test_filename $test_file --output_dir $output_dir --max_source_length $source_length --max_target_length $target_length --beam_size $beam_size --eval_batch_size $batch_size
dev_file=$data_dir/$lang/valid_ruby.jsonl
test_file=$data_dir/$lang/test_ruby.jsonl
test_model=$output_dir/checkpoint-best-bleu/pytorch_model.bin #checkpoint for test

python run.py --do_test --model_type roberta --model_name_or_path microsoft/codebert-base --load_model_path $test_model --dev_filename $dev_file --test_filename $test_file --output_dir $output_dir --max_source_length $source_length --max_target_length $target_length --beam_size $beam_size --eval_batch_size $batch_size

The above instruction for ruby. Please update the test & dev files for other languages.

GraphCodeBERT: For monolingual and multilingual fine-tuning, follow the exact instructions followed for CodeBERT, just replace the "microsoft/codebert-base" with "microsoft/graphcodebert-base". 

 

Code Search

GraphCodeBERT: for monolingual fine-tuning and inference, please clone the "CodeBERT" repo and follow the instruction in the given link

https://github.com/microsoft/CodeBERT/tree/master/GraphCodeBERT/codesearch

In the bottom please of the repo page please find the link to CodeBERT codesearch ( https://drive.google.com/file/d/1ZO-xVIzGcNE6Gz9DEg2z5mIbBv4Ft1cK/view.). 

For multilingual fine-tuning replace the run.py with the content of given run_search.py and copy the dataset "combine_retrieve" to dataset folder and run the following commands. 

lang=combine_retrieve
mkdir -p ./saved_models/$lang
python run.py \
    --output_dir=./saved_models/$lang \
    --config_name=microsoft/graphcodebert-base \
    --model_name_or_path=microsoft/graphcodebert-base \
    --tokenizer_name=microsoft/graphcodebert-base \
    --lang=$lang \
    --do_train \
    --train_data_file=dataset/$lang/train.jsonl \
    --eval_data_file=dataset/$lang/valid.jsonl \
    --test_data_file=dataset/$lang/test.jsonl \
    --codebase_file=dataset/$lang/codebase.jsonl \
    --num_train_epochs 10 \
    --code_length 256 \
    --data_flow_length 64 \
    --nl_length 128 \
    --train_batch_size 32 \
    --eval_batch_size 64 \
    --learning_rate 2e-5 \
    --seed 123456 2>&1| tee saved_models/$lang/train.log

python run.py \
    --output_dir=./saved_models/$lang \
    --config_name=microsoft/graphcodebert-base \
    --model_name_or_path=microsoft/graphcodebert-base \
    --tokenizer_name=microsoft/graphcodebert-base \
    --lang=$lang \
    --do_eval \
    --do_test \
    --train_data_file=dataset/$lang/train.jsonl \
    --eval_data_file=dataset/$lang/valid_ruby.jsonl \
    --test_data_file=dataset/$lang/test_ruby.jsonl \
    --codebase_file=dataset/$lang/codebase_ruby.jsonl \
    --num_train_epochs 10 \
    --code_length 256 \
    --data_flow_length 64 \
    --nl_length 128 \
    --train_batch_size 32 \
    --eval_batch_size 64 \
    --learning_rate 2e-5 \
    --seed 123456 2>&1| tee saved_models/$lang/test.log

The above instruction for ruby. Please update the following lines for other languages.

    --eval_data_file=dataset/$lang/valid_ruby.jsonl \
    --test_data_file=dataset/$lang/test_ruby.jsonl \
    --codebase_file=dataset/$lang/codebase_ruby.jsonl \

For multilingual CodeBERT no need to replace the run.py, just update the commands shown above to complete the task.

Extreme Summarization:

This one is identical to code summarization except the evaluation criteria. Replace the dataset and code folder in code summarization, adjust the target length to 10 and run the same set of instructions. Please find the dataset and code in "name_prediction.zip" file.  

 

PolyGlot GraphCodeBERT for Code Summarization Weights

PolyGlot CodeBERT for Code Summarization Weights


 

Files

combine.zip

Files (868.5 MB)

Name Size
md5:97b1ab653e41e994e4c710a85a0cdc82
240.8 MB Preview Download
md5:b108fb826a130ff963842aab4809920d
627.7 MB Preview Download
md5:145d1f8dae20317765cd0dd24d456b08
21.4 kB Download