Train on massive datasets without downloading anything - data streams directly from the Hub.
Teaches Qwen Latin using 1.47M texts from FineWeb-2, streamed directly from the Hub.
Blog post: Train on Massive Datasets Without Downloading
hf jobs uv run latin-llm-streaming.py
--flavor a100-large
--timeout 2h
--secrets HF_TOKEN
--… See the full description on the dataset page:
https://huggingface.co/datasets/uv-scripts/training.