LLaMA 33b finetuned on wikitext_document_level with a combination of both linear and NTK-aware ROPE scaling.
Trained with alpha=4, scale=2. Definitely works for sequence lengths up to and including 4096. Might work for much longer, but I don't have the VRAM to test properly. ¯\_(ツ)_/¯
Perplexity Graph
Training procedure
The following bitsandbytes quantization config was used during training: