A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page:
https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.