Views
No views yet
| Source weights | BF16 GGUF (unsloth/inkling-GGUF, BF16/, 41 shards, ~1.8 TiB) |
| Collection | BF16-direct (not collected on a quantized model) |
| Calibration corpus | wikitext-2-raw train, full corpus |
| Chunks | 6,400 x 512 ctx = ~3.28M tokens |
| Build | llama.cpp PR #25731 (inkling architecture), CUDA build, CPU-streamed inference |
| Hardware | 2x RTX PRO 6000 Blackwell workstation, weights streamed from PCIe 5.0 NVMe |
| Wall time | ~10.6 hours |
1llama-quantize --imatrix inkling-bf16.imatrix \
2 inkling-BF16-00001-of-00041.gguf inkling-IQ1_S.gguf IQ1_S