No auxiliary solution context is included.
Each row is a two-message conversation ending in the literal assistant target yes. Training uses the dense reward column as the soft target P(correct); correct is only the legacy boolean projection. Hidden model reasoning is not included in endpoint state.
Source: asingh15/tcs-qwen36-27b-direction-rollouts-partial-1017-mc8
Source revision: 685b614552d9ba5e2f27d6aeee1ff9d6e83c9635
Run… See the full description on the dataset page:
https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-value-partial-1017-mc8-v1-none.