English → Indian language translations of sumanthd/c3po-mixed-pretraining-corpus.
Each language is a separate dataset config (subset) on the Hub for easier browsing.
Documents are translated with Gemma 4 31B via vLLM in ~4k-token chunks at
paragraph/sentence boundaries, then stitched into full documents. Individual
chunk translations are preserved in chunks_json.
Total rows: 145,932
Source corpus:… See the full description on the dataset page:
https://huggingface.co/datasets/sumanthd/c3po-mixed-pretraining-corpus-translations.