This is a deterministic capability-preserving rewrite of
leonli66/stage3-final-mixture for LCLM Stage-3 post-training.
Only the reasoning_data and dolci_think subsets change. Their
compression_prompt is the ordinary prompt. A deterministic 50% arm keeps
the complete assistant target as ordinary SFT; the other arm wraps the inferred
reasoning prefix in <|memory_start|>...<|memory_end|> while keeping the final
answer trainable. All… See the full description on the dataset page:
https://huggingface.co/datasets/leonli66/stage3-final-mixture-cot50.