Model-level knowledge-distillation tuning was performed after layer compression.
The tuning loss decreased from 1.9244 to 1.7136 over 8 reported epochs.
Important
This file is a NanoQuant checkpoint. It is not a GGUF file and is not a
standard Transformers save_pretrained checkpoint.
It requires the NanoQuant code and compatible CUDA kernels for loading and
inference. It cannot be loaded directly with llama.cpp, Ollama, vLLM, or
AutoModelForCausalLM.from_pretrained().