Qwen3.8-27B ARA Abliterated / Uncensored NVFP4 + MTP
This is a multimodal, refusal-ablated (commonly described as "uncensored") W4A4
NVFP4 derivative of the reproducible
trohrbaugh/Qwen3.8-27B-heretic-ara
BF16 checkpoint. It is intended for native Blackwell NVFP4 inference.
What is preserved
- The vision tower, recurrent convolutions, language head, and all 15 native MTP
tensors remain BF16.
- The MTP tensors were grafted from the hash-verified source after Transformers
serialization and verified bit-exact.
- All 333 vision tensors were verified bit-exact against the BF16 source.
- The language-model linear layers use compressed-tensors NVFP4 W4A4 group-16
quantization.
Validation
This artifact passed its text-capability, image-vision, benign refusal-surface,
native MTP-acceptance, integrity, and clean-load gates on vLLM 0.23. See
BUILD_MANIFEST.json, VALIDATION_REPORT.json, and SHA256SUMS for exact
provenance and results. Video tensors/processors are preserved, but video input
was not part of the live runtime gate.
The live gate used Qwen3_5ForConditionalGeneration, native three-token MTP,
the FlashInfer CUTLASS NVFP4 kernel, an 8,192-token context, and an RTX PRO 6000
Blackwell GPU. Loaded model memory was approximately 19.53 GiB.
vLLM example
1vllm serve aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP \
2 --served-model-name Qwen3.8-27B-ARA-NVFP4-MTP \
3 --max-model-len 8192 \
4 --speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Use a recent vLLM build with Qwen3.5 multimodal and compressed-tensors NVFP4
support. Native NVFP4 execution requires compatible Blackwell hardware and CUDA
runtime support.
Notes
"Abliterated" or "uncensored" describes the source checkpoint's refusal-ablation
process; it is not a guarantee that every prompt will receive a particular answer.
Users remain responsible for evaluating outputs and applying safeguards appropriate
to their deployment.
Support
If this model is useful to you, Bitcoin donations are welcome:
bc1q5ayht3fxhj0v95fk0z8l2f6900g3awdsw5842p