Views
No views yet
| Detail | Value |
|---|---|
| Tool | spark-auto-round v0.14.3 by whpthomas |
| Format | W4A16 (INT4 weights, group_size=128, symmetric) |
| Dataset | opencode-instruct (128 samples, seqlen=2048) |
| Iterations | 200 per block |
| Batch size | 2 |
| Hardware | DGX Spark (GB10, 128 GB unified) |
| Peak RAM | 115.71 GB |
| Peak VRAM | 22.65 GB |
| Metric | Value |
|---|---|
| Avg cosine similarity | 0.9188 |
| Avg PSNR | 50.8 dB |
| Blocks | 39 transformer layers quantized |
| Format | Size |
|---|---|
| Original (BF16) | ~62.2 GB |
| Quantized (INT4) | ~20 GB (5 safetensor shards) |
quantization: auto-round, uses Marlin kernel):1vllm serve yourname/Ornith-1.0-35B-int4-AutoRound \
2 --trust-remote-code \
3 --enable-auto-tool-choice \
4 --tool-call-parser qwen3_xml \
5 --reasoning-parser qwen3 \
6 --load-format safetensors