Supervised fine-tuning data for reproducing Agent Learning via Early Experience across 8 agent environments. Each environment provides data for three training paradigms:
IL — Imitation Learning: expert
SR — Self-Reflection: expert + reflection
IWM — Implicit World Modeling: iwm (world model) → expert