Each training row contains exactly one chosen response and one rejected response. Rows are selected from the v4-real candidate pool by balancing the primary chosen action dimension and prioritizing hard/minority action buckets such as wider framing, step-back camera moves, focus, and light/exposure actions.
Images are not duplicated. Use the existing v2 image bundle on the server as IMAGE_BUNDLE.