Paper | GitHub | π€ DataFlow Collection
DataFlow is a data preparation and training system designed to parse, generate, process, and evaluate high-quality data from noisy sources (PDF, plain-text, low-quality QA), thereby improving the performance of large language models (LLMs) in specific domains through targeted training (Pre-training, Supervised Fine-tuning, RL training) or RAG using knowledge base cleaning. DataFlow has been empirically validated to⦠See the full description on the dataset page:
https://huggingface.co/datasets/OpenDCAI/dataflow-demo-Reasoning.