Views
No views yet
Standard data training, mature RL strategy, additional anti-duplicate reinforcement learning, suitable for normal use, normal output text quality, and divergent thinking in a few cases.
Incremental training of 0.4T novel content 100K SFT data generated by TifaMax, 10K SFT data generated by DeepseekR1, 2K high-quality artificial data 30K DPO reinforcement learning data generated by TifaMax to prevent duplication, enhance contextual association, and improve political security 16k ultra-long context training Random truncation training enhances robustness 8×H20 GPU full-scale fine-tuning