CodeT5 is a family of encoder-decoder language models for code from the paper:
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation by Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi.
The checkpoint included in this repository is denoted as
CodeT5-large (770M), which is introduced by the paper:
CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning by Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, Steven C.H. Hoi.
CodeT5-large was pretrained on
CodeSearchNet data in six programming languages (Ruby/JavaScript/Go/Python/Java/PHP). See Section 4.1 of the
paper for more details.
CodeT5-large was pretrained using masked span prediction objective for 150 epochs. See Section 4.1 of the
paper for more details.
We validate the effectiveness of this checkpoint pretrained with simplified strategies on
CodeXGLUE benchmark. See Appendix A.1 of the
paper for more details.
1from transformers import AutoTokenizer, T5ForConditionalGeneration
2tokenizer = AutoTokenizer.from_pretrained("Salesforce/codet5-large")
3model = T5ForConditionalGeneration.from_pretrained("Salesforce/codet5-large")
4text = "def greet(user): print(f'hello <extra_id_0>!')"
5input_ids = tokenizer(text, return_tensors="pt").input_ids
6
7# simply generate a single sequence
8generated_ids = model.generate(input_ids, max_length=8)
9print(tokenizer.decode(generated_ids[0], skip_special_tokens=True))
This release is for research purposes only in support of an academic paper. Our models, datasets, and code are not specifically designed or evaluated for all downstream purposes. We strongly recommend users evaluate and address potential concerns related to accuracy, safety, and fairness before deploying this model. We encourage users to consider the common limitations of AI, comply with applicable laws, and leverage best practices when selecting use cases, particularly for high-risk scenarios where errors or misuse could significantly impact people’s lives, rights, or safety. For further guidance on use cases, refer to our AUP and AI AUP.
1@inproceedings{CodeT52021,
2 author = {Yue Wang and Weishi Wang and Shafiq R. Joty and Steven C. H. Hoi},
3 title = {CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation},
4 booktitle = {EMNLP},
5 pages = {8696--8708},
6 publisher = {Association for Computational Linguistics},
7 year = {2021}
8}
9
10@article{CodeRL2022
11 author = {Hung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese, Steven C.H. Hoi},
12 title = {CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning},
13 journal = {arXiv preprint},
14 volume = {abs/2207.01780},
15 year = {2022}
16}