This repository contains sequentially sharded parts extracted from Google's Natural Questions dataset to optimize training and ingestion loops for LLM fine-tuning.
Dataset Structure
Format: JSON Lines (.jsonl)
Shards Uploaded: train-00000.jsonl to train-00325.jsonl (Part-1)
Data Configuration: Out-of-the-box support for datasets loader.
Generated and uploaded sequentially via RunPod pipeline.