Row-subset of the pre-tokenized trajectories in
togethercomputer/CoderForge-Preview
(trajectories-tokenized_qwencoder subset).
Size: 31,600 rows (source: 155,144 across 4 slugs).
Format: native pre-tokenized data for Qwen3 (tokenizer shared with Qwen2.5-Coder / Qwen3-Coder / Qwen3-8B).
Per row columns:
input_ids: list[int32]
attention_mask: list[int8] (all 1s; added by this subsetter so axolotl's
auto-detection of pre-tokenized datasets triggers —… See the full description on the dataset page:
https://huggingface.co/datasets/laion/CoderForge-Preview-v3-31600.