This is
larryvrh/Yi-34B-200K-Llamafied, with instruction tuning performed with Jon Durbin's
jondurbin/airoboros-3.1 dataset. That base model is
01-ai/Yi-34B-200k, but using llama2 model definitions and tokenizer to remove any remote code requirements.
The finetune was performed with 1x RTX 6000 Ada (~80 hours to this checkpoint). Prompts were truncated to 4096 tokens (for speed and VRAM headroom).
I have done very little testing with this model, so feedback on real world performance is appreciated!
Use as you would any other Hugging Face fp16 llama-2 model.
Model was trained with llama-2 chat prompt format. See
jondurbin/airoboros-l2-13b-3.1.1 model card for details.