KletterMix-12B-0.60 is the quality-filtered 0.60 release variant of KletterMix-12B, the German pretraining corpus accompanying KletterMix: Climbing Toward High-Quality German Pretraining Data.
The dataset contains German-language text examples selected with a target-language proxy score threshold of proxy_score >= 0.60. It follows the same public schema as KletterMix/KletterMix-12B and is intended for language-model pretraining, annealing experiments, data… See the full description on the dataset page:
https://huggingface.co/datasets/AIML-TUDA/KletterMix-12B-0.60.