EvoTrain-180K is a diverse dataset designed for the joint optimization of latent memory and long-context retrieval, presented in the paper EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory.
🔗 GitHub Repository | 🏠 Project Page | 📚 Paper
The dataset uses an intuitive, chat-style fine-tuning format designed for joint SFT (Supervised Fine-Tuning) and retrieval optimization. Each training instance… See the full description on the dataset page:
https://huggingface.co/datasets/MiG-NJU/EvoTrain-180K.