This repository contains a full-parameter supervised fine-tune of
Qwen/Qwen2.5-3B-Instruct
for multi-turn interaction with the WebShop environment. Improved using
Qwen.
The released weights are checkpoint global_step_160, selected because it had
the highest success rate among the evaluated SFT checkpoints on one fixed set
of 256 held-out WebShop goals. The repository also includes the 3,000
environment-verified teacher trajectories, the clean SFT messages, the exact
processed train/validation split, and checkpoint-comparison artifacts.
All models were evaluated greedily on the same 256 held-out goals with
max_turns=9, evaluation seed 123, and goal-generation seed 20260819.
Step 0 is the untouched Qwen2.5-3B-Instruct baseline.
Model
Step
Success
Mean episodic/raw return
Valid-action rate
Mean actions
Baseline
0
0.015625
0.141494
0.631061
7.7695
SFT
80
0.046875
0.290312
0.755315
6.5469
SFT (released)
160
0.054688
0.308721
0.754712
6.3398
SFT
240
0.031250
0.275714
0.731383
6.2266
SFT
320
0.050781
0.297969
0.770179
6.2734
SFT
376
0.046875
0.275796
0.792147
6.3281
The released checkpoint raised success from 4/256 to 14/256: an absolute gain
of 3.906 percentage points on this evaluation set. Mean episodic return rose
from 0.141494 to 0.308721, and valid-action rate rose by 12.365 percentage
points. These are point estimates from one fixed evaluation set; no confidence
interval or multi-seed significance claim is made.
The teacher-data pipeline used qwen3.6-35b-a3b and the train split of
webshop-small. Generation was oracle-assisted for search efficiency, but
every accepted trajectory still had to pass the environment and quality gates:
exact binary success and raw reward of 1.0;
successful episode termination and purchase;
every action valid on the current page;
an instruction/product semantic-consistency gate with minimum confidence
0.85;
one short rationale and exactly one action per assistant turn;
no validation/test goals used for teacher generation.
The resulting source contains 3,000 unique successful training seeds, 284
unique target ASINs, and an average of 5.164 interaction turns per trajectory
(range 3–9).
For training, one manually audited borderline seed (1344) was excluded,
examples were capped at 20 per target ASIN, and an ASIN-disjoint 90/10 split was
constructed:
Split statistic
Value
Quality-eligible source rows
2,999
Selected rows after ASIN cap
1,672
Train rows / unique ASINs
1,505 / 249
Validation rows / unique ASINs
167 / 35
Train-validation ASIN overlap
0
Data files
data/raw/sft_messages.jsonl: 3,000 clean multi-turn positive SFT examples.
This is the recommended human-readable SFT source.
data/raw/trajectories.jsonl: complete audit records, including environment
states, available actions, rewards, semantic-gate output, and private teacher
audit fields. Use it for auditing, not directly as the SFT input.
data/processed/train.parquet and val.parquet: the exact balanced,
ASIN-disjoint data used by the SFT run.
data/processed/split_index.json: selected example IDs and split assignment.
data/DATASET_MANIFEST.json: public generation and preprocessing metadata.
Only sft_messages.jsonl or the processed parquet files should be used as
positive SFT data. Rejected attempts and API logs are intentionally not
published.
Training configuration
Training used the official multi-turn FSDP SFT trainer from
verl, vendored through the RAGEN
experiment repository.
Setting
Value
Base model
Qwen/Qwen2.5-3B-Instruct
Training type
Full-parameter SFT
Epochs / optimizer steps
1 / 376
Learning rate
5e-6
Precision
BF16
Maximum sequence length
8,192
Effective global batch size
4
Hardware
4 × NVIDIA A100-PCIE-40GB
Checkpoint interval
80 optimizer steps, plus final step
The optimizer portion took 3,650.91 seconds (4.0566 A100 GPU-hours). The full
Slurm job, including setup, checkpoint I/O, evaluation, and plotting, used an
estimated 7.6471 allocated A100 GPU-hours. See
results/gpu_time_record.md for accounting
details.
Expected interaction format
The model was trained with the following system instruction:
You are a WebShop shopping agent. Follow the shopping instruction by interacting with the current page. At every turn, choose exactly one available action. Respond with exactly <think>brief rationale</think><answer>one action</answer> and no additional text.
Assistant turns follow this structure:
<think>I should search for the requested product category.</think><answer>search[product keywords]</answer>
The answer must be an action allowed by the current WebShop page, such as
search[...] or click[...].
Loading the model
python
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
34repo_id ="ZuoHaotong/Qwen2.5-3B-Instruct-SFT-in-WebShop"56tokenizer = AutoTokenizer.from_pretrained(repo_id)7model = AutoModelForCausalLM.from_pretrained(8 repo_id,9 torch_dtype=torch.bfloat16,10 device_map="auto",11)1213messages =[14{15"role":"system",16"content":(17"You are a WebShop shopping agent. Follow the shopping instruction "18"by interacting with the current page. At every turn, choose exactly "19"one available action. Respond with exactly <think>brief rationale"20"</think><answer>one action</answer> and no additional text."21),22},23{24"role":"user",25"content":"Shopping instruction and the current WebShop page go here.",26},27]2829text = tokenizer.apply_chat_template(30 messages,31 tokenize=False,32 add_generation_prompt=True,33)34inputs = tokenizer(text, return_tensors="pt").to(model.device)35with torch.no_grad():36 output = model.generate(**inputs, max_new_tokens=256, do_sample=False)37response = tokenizer.decode(38 output[0, inputs.input_ids.shape[1]:],39 skip_special_tokens=True,40)41print(response)
Limitations
The absolute held-out success rate is 5.47%; this is a specialized SFT
initialization for further Agent RL research, not a production shopping
system.
Checkpoint 160 was selected after comparing checkpoints on one fixed set of
256 goals. The results may contain checkpoint-selection variance and do not
establish statistical significance.
The success curve is non-monotonic, while action validity continues to
improve at later checkpoints. More SFT steps are therefore not uniformly
better for the selected task metric.
The data-generation pipeline used private oracle assistance. Oracle metadata
is retained only in the audit-rich trajectory file; clean conversation
content was checked for private-guidance leakage.
The model is specialized to the formatting and action space used by this
WebShop setup and should not be assumed to generalize to real commerce sites.
No GRPO training is included in these released weights; this is the
SFT checkpoint intended to initialize those experiments.
License and attribution
The model is a derivative of Qwen/Qwen2.5-3B-Instruct and is distributed
under the Qwen Research License Agreement, including its non-commercial-use
restriction. See LICENSE and the upstream
license page.
The repository's previous Apache-2.0 placeholder did not match the license
shipped with the local upstream 3B checkpoint and has been corrected.
Qwen is licensed under the Qwen Research License Agreement, Copyright (c)
Alibaba Cloud. All Rights Reserved. This repository contains modified model
weights and must not be interpreted as an official Qwen release.
The included WebShop-derived data and evaluation artifacts are provided for
research reproducibility. Users are responsible for complying with the terms
of the upstream WebShop resources and any applicable dataset restrictions.