Views
No views yet
ibm-granite/granite-4.0-1b.granite-4.0-1b-cn-f32.gguf: merged F32 GGUF exported from the local MLX LoRA run.granite-4.0-1b-cn-q8_0.gguf: merged Q8_0 GGUF quantized from the merged model.mlx/output/granite_4_0_1b_tc_trs_1epoch_alltargets_8k_bs2/mlx/data/granite_4_0_1b_tc_trs/rank=8, scale=16, batch_size=2, iters=530, max_seq_length=8192llama-export-loraQ8_0 with llama-quantizellama.cpp CLI smoke tests succeeded for both merged artifacts.llama.cpp.@@@@@@@@... with empty tool_calls for:
Q8_0 GGUFQ8_0 GGUFF32 GGUFunsloth/granite-4.0-1b q8 GGUFllama.cpp.llama.cpp path using Granite's own tokenizer chat template and XML tool-call parsing fixes.llama-cpp-python and LM Studio runtime checks diverged and should not be treated as equivalent validated Granite runtimes.docs/experiments.md and docs/granite_issues.md in the training repo for the full debugging and runtime findings.