Views
No views yet
A number of RL strategies are used, mainly using 671B R1 distilled data, with high output divergence, inheriting the advantages of R1, and also inheriting the harmfulness of R1. Good literary performance.
Incremental training of 0.4T novel content 40K SFT data generated by TifaMax, 60K SFT data generated by DeepseekR1, 2K high-quality artificial data 30K DPO reinforcement learning data generated by TifaMax to prevent duplication, enhance context association, and improve political security 10K PPO data generated by TifaMax, 10K PPO data generated by DeepseekR1 16k ultra-long context training Random truncation training enhances robustness 8×H20 GPU full-scale fine-tuning