Views
No views yet
| Component | Method | Weight format | Effective targets |
|---|---|---|---|
Decoder Linear layers and lm_head | GPTQ | symmetric INT4, group size 256, static act-order | 253 |
| Vision encoder Linear layers | GPTQ | symmetric INT8, per-channel, static act-order | 162 |
| Projector/architecture exceptions | none | source BF16 | Explicitly recorded in recipe.yaml |
lm_head follows the default GPTQ INT4/G256 policy; no INT8 escalation was made. Calibration used lmms-lab/flickr30k test[:512], deterministic dataset order, 512 image-text samples, sequence length 2048, batch size 1, no concatenation, and no padding to maximum length. The exact target names and effective settings are saved in recipe.yaml.7055d3010c38ccb5dcae1bc9535ca19c7fe5d79f, deterministic generation, and exact_matching without an API judge.| Checkpoint | Overall accuracy |
|---|---|
| Source | 54.1333% |
| Quantized | 53.5333% |
| Absolute change | -0.6000 percentage points |
| Relative change | -1.1084% |
vlmeval_results/mmstar_results.csv; raw predictions and status files remain under vlmeval_results/mmstar_full/.EleutherAI/wikitext_document_level, wikitext-2-raw-v1, no example limit; 2026-07-18): source 13.8655; this checkpoint 15.5220; absolute degradation +1.6565; relative degradation +11.95%. Raw JSON evidence: perplexity_results/vlm_full_wikitext/qwen25_vl_3b_{source,gptq_w4g256_visiongptq_w8}_v3.json.quantize.py, recipe.yaml, and versions.txt used for the artifact:1python quantize.py \
2 --model-path /path/to/source-model \
3 --output-dir /path/to/new-quantized-model1Python: 3.13.2
2torch: 2.11.0+cu128
3CUDA: 12.8
4transformers: 5.10.1
5llm-compressor: 0.12.0
6auto-round: 0.13.0
7compressed-tensors: 0.17.1recipe.yaml.