GGUF is a new format introduced by the llama.cpp team on August 21st 2023. It is a replacement for GGML, which is no longer supported by llama.cpp.
Here is an incomplete list of clients and libraries that are known to support GGUF:
🙏 Special thanks to
Georgi Gerganov and the whole team working on
llama.cpp for making all of this possible.
FireFunction is a state-of-the-art function calling model with a commercially viable license. View detailed info in our
announcement blog. Key info and highlights:
🔆 Support of parallel function calling (unlike FireFunction v1) and good instruction following
💡 Hosted on the
Fireworks platform at < 10% of the cost of GPT 4o and 2x the speed