The ~24.5B-token pretraining corpus used to train Tessera 1B, from AIIT-THRESHOLD.
What makes it different: this is not a single lab dump. The bulk is web text (DCLM), but the long tail was hand-selected, source by source — a deliberate curation rather than a firehose.
Shards
674, tokenized with the Tessera tokenizer (byte-level BPE, vocab 65,536, memory-organ… See the full description on the dataset page:
https://huggingface.co/datasets/AIIT-Threshold/AIIT-Tessera24B-dataset.