Multilingual training for Software Engineering
Authors/Creators
Description
This is the replication package for "Multilingual training for Software Engineering".
Code Summarization
CodeBERT: for monolingual fine-tuning and inference, please clone the "CodeXGLUE" repo and follow the instruction in the given link
https://github.com/microsoft/CodeXGLUE/tree/main/Code-Text/code-to-text
For multilingual fine-tuning and inference, copy the combine.zip folder to dataset folder and unzip it. Run the following commands:
cd code
lang=combine #programming language
lr=5e-5
batch_size=32
beam_size=10
source_length=256
target_length=128
data_dir=../dataset
output_dir=model/$lang
train_file=$data_dir/$lang/train.jsonl
dev_file=$data_dir/$lang/valid.jsonl
epochs=10
pretrained_model=microsoft/codebert-base #Roberta: roberta-base
python run.py --do_train --do_eval --model_type roberta --model_name_or_path $pretrained_model --train_filename $train_file --dev_filename $dev_file --output_dir $output_dir --max_source_length $source_length --max_target_length $target_length --beam_size $beam_size --train_batch_size $batch_size --eval_batch_size $batch_size --learning_rate $lr --num_train_epochs $epochs
batch_size=64
dev_file=$data_dir/$lang/valid.jsonl
test_file=$data_dir/$lang/test.jsonl
test_model=$output_dir/checkpoint-best-bleu/pytorch_model.bin #checkpoint for test
python run.py --do_test --model_type roberta --model_name_or_path microsoft/codebert-base --load_model_path $test_model --dev_filename $dev_file --test_filename $test_file --output_dir $output_dir --max_source_length $source_length --max_target_length $target_length --beam_size $beam_size --eval_batch_size $batch_size
dev_file=$data_dir/$lang/valid_ruby.jsonl
test_file=$data_dir/$lang/test_ruby.jsonl
test_model=$output_dir/checkpoint-best-bleu/pytorch_model.bin #checkpoint for test
python run.py --do_test --model_type roberta --model_name_or_path microsoft/codebert-base --load_model_path $test_model --dev_filename $dev_file --test_filename $test_file --output_dir $output_dir --max_source_length $source_length --max_target_length $target_length --beam_size $beam_size --eval_batch_size $batch_size
The above instruction for ruby. Please update the test & dev files for other languages.
GraphCodeBERT: For monolingual and multilingual fine-tuning, follow the exact instructions followed for CodeBERT, just replace the "microsoft/codebert-base" with "microsoft/graphcodebert-base".
Code Search
GraphCodeBERT: for monolingual fine-tuning and inference, please clone the "CodeBERT" repo and follow the instruction in the given link
https://github.com/microsoft/CodeBERT/tree/master/GraphCodeBERT/codesearch
In the bottom please of the repo page please find the link to CodeBERT codesearch ( https://drive.google.com/file/d/1ZO-xVIzGcNE6Gz9DEg2z5mIbBv4Ft1cK/view.).
For multilingual fine-tuning replace the run.py with the content of given run_search.py and copy the dataset "combine_retrieve" to dataset folder and run the following commands.
lang=combine_retrieve
mkdir -p ./saved_models/$lang
python run.py \
--output_dir=./saved_models/$lang \
--config_name=microsoft/graphcodebert-base \
--model_name_or_path=microsoft/graphcodebert-base \
--tokenizer_name=microsoft/graphcodebert-base \
--lang=$lang \
--do_train \
--train_data_file=dataset/$lang/train.jsonl \
--eval_data_file=dataset/$lang/valid.jsonl \
--test_data_file=dataset/$lang/test.jsonl \
--codebase_file=dataset/$lang/codebase.jsonl \
--num_train_epochs 10 \
--code_length 256 \
--data_flow_length 64 \
--nl_length 128 \
--train_batch_size 32 \
--eval_batch_size 64 \
--learning_rate 2e-5 \
--seed 123456 2>&1| tee saved_models/$lang/train.log
python run.py \
--output_dir=./saved_models/$lang \
--config_name=microsoft/graphcodebert-base \
--model_name_or_path=microsoft/graphcodebert-base \
--tokenizer_name=microsoft/graphcodebert-base \
--lang=$lang \
--do_eval \
--do_test \
--train_data_file=dataset/$lang/train.jsonl \
--eval_data_file=dataset/$lang/valid_ruby.jsonl \
--test_data_file=dataset/$lang/test_ruby.jsonl \
--codebase_file=dataset/$lang/codebase_ruby.jsonl \
--num_train_epochs 10 \
--code_length 256 \
--data_flow_length 64 \
--nl_length 128 \
--train_batch_size 32 \
--eval_batch_size 64 \
--learning_rate 2e-5 \
--seed 123456 2>&1| tee saved_models/$lang/test.log
The above instruction for ruby. Please update the following lines for other languages.
--eval_data_file=dataset/$lang/valid_ruby.jsonl \
--test_data_file=dataset/$lang/test_ruby.jsonl \
--codebase_file=dataset/$lang/codebase_ruby.jsonl \
For multilingual CodeBERT no need to replace the run.py, just update the commands shown above to complete the task.
Extreme Summarization:
This one is identical to code summarization except the evaluation criteria. Replace the dataset and code folder in code summarization, adjust the target length to 10 and run the same set of instructions. Please find the dataset and code in "name_prediction.zip" file.
PolyGlot GraphCodeBERT for Code Summarization Weights
PolyGlot CodeBERT for Code Summarization Weights