Views
No views yet
| Parameter | Value |
|---|---|
| Method | AutoRound (W4A16) |
| Group Size | 16 |
| Symmetric | Yes |
| Iterations | 1000 |
| Calibration Samples | 512 |
| Sequence Length | 4096 |
| Torch Compile | Enabled |
quant_nontext_module): False (Kept in BF16 to preserve visual reasoning and OCR precision)layer_config): Multi-Token Prediction (mtp, mtp.fc) kept in native bfloat16--speculative_config '{"method":"mtp","num_speculative_tokens":1}'num_speculative_tokens=1 is a stable default for balancing speed and accuracy.transformers and backends that support AutoRound GPTQ-format weights (e.g., vLLM, SGLang, AutoGPTQ). For full model details, architecture, and capabilities, refer to the base model page.🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
| Template | CUDA Version | Docker Image | Template ID | Deploy |
|---|---|---|---|---|
| PyTorch 2.13 (CUDA 12.6) | 12.6 | vishva123/cuda-12.6-pytorch-2.13-runpod | gmlupxnxfk | |
| PyTorch 2.13 (CUDA 13.0) | 13.0 | vishva123/cuda-13.0-pytorch-2.13-runpod | y3j8xvk4f4 | |
| PyTorch 2.13 (CUDA 13.2) | 13.2 | vishva123/cuda-13.2-pytorch-2.13-runpod | vigpissn5w |
| Template | CUDA Version | Docker Image | Template ID | Deploy |
|---|---|---|---|---|
| PyTorch 2.12 (CUDA 12.6) | 12.6 | vishva123/cuda-12.6-pytorch-2.12-runpod | ctmz86zmf0 | |
| PyTorch 2.12 (CUDA 13.0) | 13.0 | vishva123/cuda-13.0-pytorch-2.12-runpod | qjko5yiwzi | |
| PyTorch 2.12 (CUDA 13.2) | 13.2 | vishva123/cuda-13.2-pytorch-2.12-runpod | ifg6xmye0f |