Views
No views yet
Qwen/Qwen3-4B-MLX-4bit trained for strict local
Hermes-style tool-call output.https://huggingface.co/edithatogo/qwen3-4b-hermes-lora/no_think<think>\n\n</think>\n\nQwen/Qwen3-4B-MLX-4bitgemma4/scripts/train_config.qwen3-4b.strict-toolcall-v4-targeted.yamlgemma4/data/strict_tool_call/expanded_splits_v4_targetedgemma4/experiments/qwen3-4b-strict-toolcall-v4-targeted/lora_adapterreports/publication/qwen3-4b-strict-toolcall-v4-targeted/dataset-token-audit.jsonreports/publication/qwen3-4b-strict-toolcall-v4-targeted/dataset-overlap-audit.json| Suite | Pass | JSON valid | Arguments | Invalid tool | Multi-turn |
|---|---|---|---|---|---|
benchmarks/tool_call_local/heldout_suite.json | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| Suite | Pass |
|---|---|
benchmarks/tool_call_local/suite.json | 1.000 |
| Pilot | Pass | Notes |
|---|---|---|
| BFCL-style pilot | 0.667 | local pilot only, not official BFCL |
| IFEval-style pilot | 0.667 | local pilot only, not official IFEval |
| Coding sanity pilot | 1.000 | local pilot only, not HumanEval/MBPP |
1source scripts/env.sh
2PYTHONPATH=scripts ./.venv/bin/python scripts/run_tool_call_benchmark.py \
3 --model Qwen/Qwen3-4B-MLX-4bit \
4 --adapter gemma4/experiments/qwen3-4b-strict-toolcall-v4-targeted/lora_adapter \
5 --suite benchmarks/tool_call_local/heldout_suite.json \
6 --user-prefix /no_think \
7 --assistant-prefill $'<think>\n\n</think>\n\n' \
8 --run-id qwen3-4b-strict-toolcall-v4-targeted-heldout-prefill-20260525 \
9 --max-tokens 256/Volumes/PortableSSD/hermes-evals/tool-call-benchmark/qwen3-4b-strict-toolcall-v4-targeted-heldout-prefill-20260525RUNTIME_PROMPT_PROFILES.yaml as qwen3-no-think-assistant-prefill.notify_care_team.release-decision.md; the publication
bundle is expected to pass scripts/validate_publication_bundle.py --require-ready.