Synthetic texts about astronautics and space mission engineering, in English and French.
The dataset ships as two subsets:
Subset
Rows
Tokens
Description
default
23,800
13.14M
Full set
deduplicated
13,525
4.49M
Near-duplicates removed
from datasets import load_dataset
full = load_dataset("patrickfleith/astro_texts_dataset", split="train")
dedup = load_dataset("patrickfleith/astro_texts_dataset", "deduplicated", split="train")