This repository contains the complete artifacts from one pair of 200-step
BabyAI training runs comparing RL-only against always-on ECHO 1.0.
Model: Qwen/Qwen3.5-9B
Trainer: Prime-RL v0.7.0 (d334ea52)
Verifiers: v0.2.0
Environment: agentboard-babyai-v1-context
Training split: 84 tasks (three examples per BabyAI subtask)
Held-out split: 28 tasks (one example per subtask)
Training steps: 200
Batch size: 128
Group… See the full description on the dataset page:
https://huggingface.co/datasets/bhoy/babyai-qwen35-9b-v070-alwayson-200-run1.