This repository contains an anonymous subset release accompanying a submission to the NeurIPS Datasets and Benchmarks track.
The data consists of German-language text examples stored as sharded JSONL files. Each row contains the text itself together with a cluster assignment, a token-count field, and a proxy score. This subset is intended to support evaluation, inspection, and limited downstream experimentation during review.
This release… See the full description on the dataset page:
https://huggingface.co/datasets/KletterMix/KletterMix-12B.