Views
No views yet
nvidia/GLM-5.2-NVFP4@aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa — no cumulative merge from an earlier SuperGLM checkpoint.o_proj matrices at strength 2.down_proj matrices at strength 0.15.lm_head remains byte-identical to NVIDIA's base.8e1604d1171fe8a48d8febecd22f600e462bdcdd. The standard-text development gates reuse previously opened official HarmBench partitions for candidate selection. The sealed final instead uses previously unused FunctionalCategory=contextual official-test behaviors, formatted byte-for-byte like HarmBench DirectRequest (ContextString, separator, then Behavior) and judged with the official contextual classifier template.| Gate | Physical prompts | Greedy repeats | Official no / invalid / total | Explicit refusal strings | Broad lexical flags | Quality checks | Structural anomalies | Unicode replacements |
|---|---|---|---|---|---|---|---|---|
| Runtime adapter screen | 41 | 2 | 1 / 0 / 82 | 0 | 19 | 16/16 | 0 | 0 |
| Serialized fused gate | 35 | 2 | 4 / 0 / 70 | 0 | 24 | 15/16 | 0 | 0 |
| Sealed serialized final | 80 | 2 | 2 / 0 / 160 | 0 | 24 | 16/16 | 0 | 0 |
enable_thinking=false, no explicit system prompt, and adaptive_directness=false. Runtime LoRA is used only for the first candidate screen. The two release gates run on serialized weights with adapters disabled. Acceptance requires at least 15/16 quality checks, zero structural anomalies, zero Unicode replacement characters, and zero explicit refusal strings. The official classifier release threshold is configured as fused <= 6 and sealed <= 6 refused/non-exhibiting hits, with invalid outputs fixed at zero. This v2 card reports the exact no / invalid / total counts above so the configured release slack is auditable. In HarmBench semantics a yes means that the requested behavior is present, so treating every no as refused/non-exhibiting is stricter than a refusal-string check. Broad lexical flags are reported for continuity but are not a release gate because they also match compliant caveats containing words such as “illegal” or “dangerous.”lm_head repair and is not reported as an independent benchmark. It is excluded from every official acceptance partition and is also supplied to the overlap audit.flashinfer_cutlass MoE runner on 8x RTX PRO 6000 Blackwell GPUs.1import sglang as sgl
2
3engine = sgl.Engine(
4 model_path="Jiunsong/SuperGLM-5.2-abliterated-NVFP4",
5 tp_size=8,
6 quantization="modelopt_fp4",
7 moe_runner_backend="flashinfer_cutlass",
8 disable_shared_experts_fusion=True,
9)aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aabroad62-r4-s2-obliteratus_only-shared-r2-s0p15sha256:d7ded1c1eea29006823f7219a35c71a4cfdcdf4038aaa39c4909e57626f1300ezai-org/GLM-5.2-FP8@ba978f7d347eaf65d22f1a86833408afdb953541sha256:956fd23f7a45308878fc41a55e30f8995d515ad3b53bc1b1011d35efeb021101cais/HarmBench-Llama-2-13b-cls@bda705349d1144fa618770bea64d99ce54e3835b5d6bb9e3cf4d1e5f3f7620093113222ee115af75b0b2913f00e4c7225ec9f219788f4f6aa1491c433c4da76c9140cfc30966cea3ff3875c4d0fcb336d92f60e0172dc74a35e1752df75ecfb2b2cf9326d2852bb1379868ebeec9571654489679