phi-4 — W4A16 (compressed-tensors)
Standard W4A16 quantization of
microsoft/phi-4, produced with
llm-compressor (the
official vLLM-team quantization toolkit) inside a reproducible Docker
container. The artifact saves in
compressed-tensors format and
is
drop-in loadable by vLLM — no upstream patches, no client-side
shims; vLLM auto-detects the quantization config from the embedded
config.json at load time.
This release is part of an ongoing series of vLLM-friendly quantized
packs maintained by
atlas, a self-evolving agent project run by
Alex Adamopoulos at
assert.gr.
Reproducibility
| Parameter | Value |
|---|
| Source model | microsoft/phi-4 |
| Quantization tool | llm-compressor 0.12.0 (Neural Magic / vLLM team) |
| Quantization recipe | GPTQModifier |
scheme | W4A16 |
targets | Linear |
ignore | re:.*lm_head, re:.*vision_tower.*, re:.*multi_modal_projector.* |
graft (kept in source dtype) | — |
sequential_targets | — |
dampening_frac | 0.01 |
| Calibration dataset | ultrachat-200k |
| Calibration samples | 256 |
max_seq_length | 2048 |
| Quantized size | 8.47 GiB |
| Quantization time | — |
License
Inherits the license of the base model. By using this artifact you
agree to the original license at the source link above. Atlas /
assert.gr adds no additional restrictions on the quantized weights.
Usage with vLLM
1docker run --runtime=nvidia --gpus all \
2 -p 8000:8000 \
3 -e HF_TOKEN=hf_XXX \
4 vllm/vllm-openai:latest \
5 --model aleada/Phi-4-W4A16 \
6 --gpu-memory-utilization 0.92 \
7 --enable-prefix-caching
vLLM auto-detects compressed-tensors from the model's config — no
--quantization flag required (it is accepted as a redundant hint).
vLLM also picks the model's full native context window from
config.json. If you hit KV-cache OOM on a smaller GPU, pin a shorter
window with --max-model-len 16384 (or smaller) — leave it off to get
the maximum the model was trained for.
Once vLLM is running, hit it with any OpenAI client:
1from openai import OpenAI
2client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
3resp = client.chat.completions.create(
4 model="aleada/Phi-4-W4A16",
5 messages=[{"role": "user", "content": "Hello"}],
6)
7print(resp.choices[0].message.content)
Hardware target
Requires CUDA compute-capability ≥ 8.0 (Ampere or newer). Verified on
NVIDIA RTX 3090 (compute 8.6) where the W4A16 path runs the
language tower at INT4 weights / BF16 activations through vLLM's
compressed-tensors kernels.
Weight-only INT4 is the point on this class of card: FP8 and NVFP4
checkpoints are native on Hopper and Blackwell but emulated or
unusable on Ampere, where the INT4 Marlin kernels are what actually
run fast.
Check this pack yourself
Quantization can drop or disable part of a model without failing: the
pack loads, serves, and answers correctly while something its card
says it kept is absent, or present and ignored by the runtime. Nothing
errors, and the card still promises it.
Pack integrity check
reads any published repo's metadata — safetensors headers and
config.json, no weights — and reports whether its exclusion entries
name real modules, whether anything from the source model failed to
reach it, and whether anything is left at source precision without
being declared. It runs entirely in your browser, so it reads exactly
what you could read yourself.
Point it at this pack. Point it at someone else's.
About the maintainer
Alex Adamopoulos is the founder of
assert.gr and
the engineer behind the
atlas self-evolving AI agent platform.
Atlas runs a planner→executor→supervisor loop over a skill registry,
backed by Postgres, Redis, Qdrant, and a multi-LLM vLLM deployment.
Quantization releases like this one keep the open-source model
ecosystem usable on consumer-grade hardware for self-hosted agent
research.
Connect:
- HuggingFace: @aleada
- Company: assert.gr