Full-parameter SFT of
Qwen2.5-14B-Instruct for the
AppWorld interactive-coding agent benchmark, trained on
504 skill-augmented full-conversation trajectories (87 train tasks × 6
sampling temperatures, hole-repaired to near-complete grid).
Reference points under the same protocol: SAGE-14B baseline 35.7% TGC;
a 32B sibling of this recipe reaches 61.9% TGC.
The model expects the Skill-HELIX prompt contract (supervisor header, function
policy, retrieved-experience block, scenario position). It is a research
artifact, not a general assistant.
Training data was teacher-distilled on AppWorld train scenarios only;
train/dev/test scenario splits are disjoint; no evaluator internals or ground
truth appear in prompts. Full run records, manifests and fingerprints are
archived by the authors.