10% difficult-advice / 90% TULU3 replay, by token. Built for training with
loss on assistant tokens only.
mixture.jsonl is byte-identical (md5 af628722652f05debf5cffd44db09f88, 2,257 rows) to the mixture used by the
full-token arm
…-tulu-lora-10-90, so the loss mask is the only
difference between the two runs.