Views
No views yet
494c870).Fine-tuned starcoder2-15b with an additional 0.7 billion high-quality, code-related tokens for 3 epochs. We used DeepSpeed ZeRO 3 and Flash Attention 2 to accelerate the training process. It achieves 77.4 pass@1 on HumanEval-Python. This model operates using the Alpaca instruction format (excluding the system prompt).
| Layers | Context | Template |
|---|---|---|
40 | 16384 | ### Instruction {instruction} ### Response {response} |