Потом, если получится хоть 1 рабочий конверт, засуну результаты самого gguf'а
Getting started
You can use LM Studio or llama.cpp for example. In LM Studio search for model and download the Q4 version, then write some dialog, LLM will continue better if dialog is not too tiny (need 50-80+ words).
Origin
This is Llama.cpp compatible versions of an original 1.7B model + some tokenizer patches, prompt templates.