Processed evaluation datasets for Latent-SFT trajectory generation and model diagnostics.
These files are packaged for batch CoT trajectory generation. The prompt should put the final boxed answer only in the generated response / cot_answer; the reasoning-only part should not contain an extra boxed answer.
mmlu_pro_validation_audit.jsonl
70
answer… See the full description on the dataset page:
https://huggingface.co/datasets/liaialley/latent-sft-eval-benchmarks.