German: "treasure trove." 884 clean novel chapters
(2,533,296 words) from 25 US-public-domain works, harvested
from Project Gutenberg and gated entirely by
deterministic quality filters — no LLM touched the text or the admission
decisions.
Built as the human-written chosen/seed side for preference datasets
(e.g. gutenberg2-dpo-style
chapter DPO, or single-axis voice-degradation DPO where a model only ever
writes the rejected side). Useful anywhere you want… See the full description on the dataset page:
https://huggingface.co/datasets/schneewolflabs/fundgrube.