Views
No views yet
Qwen/Qwen3.5-2B after one epoch of
supervised fine-tuning on successful ShopSimulator teacher trajectories.
Training used the Pi agent harness and the modified Slime integration.| Item | Value |
|---|---|
| Base model | Qwen/Qwen3.5-2B |
| Teacher | deepseek-v4-flash, thinking disabled |
| Collection | 512 tasks, one candidate per task |
| Accepted trajectories | 412 (80.47% task coverage) |
| Turn-level SFT examples | 6,153 |
| Epochs | 1 |
| Optimizer steps | 2,051 |
| Global batch size | 3 |
| Learning rate | 1e-5 |
| Maximum tokens per GPU | 12,288 |
| Longest prepared example | 8,162 tokens at conversion time |
| Loss mask | Qwen3.5 multi-turn mask; template-injected empty think blocks excluded |
| Training framework | Slime + Megatron-LM |
| Hardware | One NVIDIA Pro 6000D 84 GB |
| Weight identity SHA-256 | 8a98272246a3dbd4c59d64e3981807e6b7719007880cf2aed9fe0996e76f8a7d |
official_test_200 task slice, and one rollout per
task. These are k=1 point estimates, not uncertainty estimates.| Model | Positive-reward pass@1 | Strict-success pass@1 | mean@1 r_loose | mean@1 r_hard |
|---|---|---|---|---|
| Qwen3.5-2B | 2.0% | 0.0% | 0.004286 | 0.000000 |
| This SFT checkpoint | 72.5% | 10.5% | 0.389829 | 0.124417 |
Qwen/Qwen3.5-2B with a
recent Transformers version that supports Qwen3.5. Refer to the
official Qwen3.5-2B model card for
loading and inference examples, replacing the model ID with this repository.pi-slime-shopsimulator source
repository when it is published.LICENSE. This model card identifies
the checkpoint as modified and does not grant rights to third-party
ShopSimulator task or environment content.