Qwen3.5-4B Clawd-RIFT — GGUF quantizations
File Size Quant clawd-rift-f16.gguf7.9 GB bf16->f16 baseline clawd-rift-Q8_0.gguf4.2 GB 8-bit (near-lossless) clawd-rift-Q5_K_M.gguf2.9 GB 5-bit K-quant (recommended balance) clawd-rift-Q4_K_M.gguf2.6 GB 4-bit K-quant (smallest practical)
Usage with llama.cpp
1 llama-cli -m clawd-rift-Q5_K_M.gguf -p 'Your prompt'
2 llama-server -m clawd-rift-Q5_K_M.gguf --port 8080
In Ollama / LM Studio: import GGUF directly. Set chat template to Hermes-style with <tool_call>{json}</tool_call> for tool use.
Evaluation results
tbench-2 (89 docker tasks via Pi-style runner)
7/89 (7.9%). Tasks unique to clawd-rift: fix-ocaml-gc, pytorch-model-recovery.
Variant in pipeline Pass on tbench-2 ckpt600 (Soyuz SFT only) 7 clawd-100 (+ ClawGym 100 steps) 7 clawd-200 (+ ClawGym 200 steps) 7 clawd-rft (positive-only SFT on rollouts) 6 clawd-rift (true RIFT on rollouts) — this model 7
ClawGym-Bench (200 tasks via openclaw scaffold)
Stat Value mean 0.371 half+ (≥0.5) 80/200 (40%) perfect (=1.0) 2 (tasks 78, 148) zero 40
Comparison to RUC-AIBOX ClawGym leaderboard (compact open-weight models):
Model ClawGym avg Qwen3-32B 33.11 Qwen3-8B 35.02 clawd-rift (this, 4B, QLoRA, 1 GPU) 37.10 Qwen3-30A3B (MoE) 45.11 ClawGym-4B (RUC-AIBOX full SFT) 47.73
Optimal inference parameters
Sampling sweet-spot is scaffold-dependent.
Scaffold Task type Optimal sampling openclaw (ClawGym-style formal spec)JSON/Markdown to schema T=0.3-0.5, top_p=0.95, no min_ppi-agent (terminus_runner shell explore)trial-and-error commands T=0.7-0.8, top_p=0.95, min_p=0.05
Universal default that loses only ~5% on each:
temperature=0.5, top_p=0.95, top_k=40, repetition_penalty=1.05
Training methodology — pipeline of 3 stages
Stage 1: Soyuz SFT (ckpt600 — base agent format)
QLoRA r=64 alpha=128 on Qwen/Qwen3.5-4B.
Datasets: AlexWortega/Soyuz-sft + AlexWortega/AgentTrove
Format: Hermes-style JSON tool calls (<tool_call>{"name":...,"arguments":...}</tool_call>)
600 steps total, seq=8K, Muon optimizer for LoRA matrices
Output: ckpt-400, ckpt-600 (intermediate); soup_sum = ckpt400 + ckpt600 (arithmetic merge)
Stage 2: ClawGym continue-train (clawd-100, clawd-200 — openclaw scaffold adaptation)
Continue-train ckpt600 on filtered
RUC-AIBOX/ClawGym-Trajectory.
1937 trajectories (filtered ≤16K tokens out of 24.5K)
200 steps, seq=16K, LR=1e-4, AdamW
Hermes chat template + openclaw native tools (read/write/exec/web_search/...)
Output: clawd-100 (mid), clawd-200 (final)
Stage 3: RIFT — own rollouts + reward feedback
True RIFT loss on top of clawd-200:
1 # positive (reward > 0): NLL × reward — weighted SFT
2 # negative (reward = 0): exp(logp) × negative_scale — unlikelihood
61 trajectories from soup_sum's own ClawGym rollouts (46 pos + 15 neg, reward 0-1)
5 epochs / 80 steps, LR=2e-5
Implementation: compare_offlinegpro/src/trainers/offline_losses.py
Output: clawd-rift ← this model
Repos
GGUF breakdown:
clawd-rift-f16.gguf (7.9 GB, baseline)
clawd-rift-Q8_0.gguf (4.2 GB, near-lossless)
clawd-rift-Q5_K_M.gguf (2.9 GB, recommended)
clawd-rift-Q4_K_M.gguf (2.6 GB, smallest)
Related: Stage-1-only model (qwen35-4b-soyuz)
A cleaner, stronger reference for the Stage-1 base (Soyuz SFT only — no ClawGym, no RIFT) is now available, trained as full bf16 LoRA r=128 (vs QLoRA r=64 here):
Final eval on Soyuz-clean held-out:
loss=0.247, token_acc=0.936. Trained on the cleaned 11-stream subset of
AlexWortega/Soyuz-sft at seq=16K, 1 epoch.
Useful if you want only the Hermes-tool-call SFT without the ClawGym/RIFT specialization.
Loading caveat
These GGUF files were converted by an older
llama.cpp build before upstream support for the Qwen3.5 hybrid linear+full attention architecture stabilized. Some llama.cpp builds may complain about
missing tensor or
unsupported architecture when loading. The merged HF weights at
qwen35-4b-clawd-rift-merged are the canonical reference.