miniVERL Qwen3-0.6B tool-policy SFT checkpoint
This frozen LoRA adapter is the common post-SFT starting point and qualified
teacher candidate for miniVERL Alignment Lab v1. It was trained on deterministic,
in-memory sandbox tasks covering authorization, confirmation, instruction
hierarchy, secret exclusion, benign completion, and safe error recovery.
The candidate was selected on the 24-task eval split only. It achieved 24/24
strict task success, 100% parse-valid tool calls, and 100% final-answer format
validity. No final-test task was read before the Alignment Lab preregistration.
Provenance:
- miniVERL source commit:
edd4b6ef542c1e07a96f61b8aba52205c23522c6
- base revision:
c1899de289a04d12100db370d81485cdf75e47ca
- source checkpoint digest:
480d3999ce31d4a7ae545d1f7a524474077460127c539f85c3040f91efc60994
- adapter config SHA-256:
a0a3d8ba706fc7de3d19434af509df2637b3ebc00c557e3341a80b568a54228a
- adapter weights SHA-256:
8765cdf50ce044264ae42f11381aba35c69f6c4ab2d71c163862eead25fceb73
- hardware: one NVIDIA GeForce RTX 4080 16 GB
The exact machine-readable provenance and eval record are in
miniverl_adapter_manifest.json.
Limitations
This is a small synthetic policy suite, not evidence of broad safety or general
alignment. All actions are sandboxed and harmless. The eval split is suitable
for candidate selection, not a final claim. Use the frozen Alignment Lab test
artifacts for method comparisons.