Partial rows include 3 same-problem endpoint solution references; reference reward labels are included. Endpoint rows use the explicit no-context sentinel.
Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state.
Identity… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-solution-label.