This repo contains GGUF-format weights for
Vocaela-2-500M-1024R2, containing:
Note: So far LlamaCpp has problem on rendering chat template correctly. To workaround it, we apply chat template (e.g., using python / node.js etc.) before calling llama-server endpoint. For examples of how to use it, please refer to repo
vocaela-500m-demo