Views
No views yet
AlexWortega/qwen35-4b-soyuz-merged.| field | value |
|---|---|
| Method | phase2 exp4 counterfactual injection (wrong-action vs right-action), mean diff L=5, strength=0.5 |
| tbench-2 (17) | 2/17 |
| HermesAgent-20 | 9 / 20 |
| MMLU-Pro | — |
| EQbench3 | — |
| Notes | MMLU not measured (initial bench lost to GPU clash, re-bench captured HA20=9/20) |
capability-vectors
sweep. Phase 1 best was v2 (HA20 8/20, MMLU collapse 58→2). Phase 2 explores
multi-token / hard-pairs / counterfactual / agent-only / activation-steering recipes.phase2/results/all_variants.csv for the live results table.1python -m sglang.launch_server \
2 --model-path AlexWortega/qwen35-4b-soyuz-abliterated-v8_cfact \
3 --dtype bfloat16 --trust-remote-code \
4 --tool-call-parser hermes --chat-template hermes_qwen.jinja