Precomputed float32 pooling embeddings derived from neulab/agent-data-collection.
The selection takes up to the first 500 standardized trajectories from each of 23 domains.
Model: Qwen/Qwen3-8B
PRM: state/action pairs, 4,096-token head truncation
TRM: full trajectories, 32,768-token tail truncation
Hard negatives: none
Reward labels: preserved when ADP details includes one; otherwise marked as assumed demonstration success