Views
No views yet

int4_awq) with FP8 KV cache, produced with NVIDIA TensorRT Model Optimizer (PTQ) from GestaltLabs/Ornstein-3.5-9B-V2 — the reinforcement-learning post-training (V2) of Ornstein 3.5 9B. Text weights only; pair with the full base model for the vision tower.hf_quant_config.json and are intended for ModelOpt-aware runtimes (TensorRT-LLM / vLLM). Load the directory as a standard HF model id with the matching backend.