FP8 training pipeline for Spider-FLEXITOKENS on NVIDIA Blackwell GPUs (sm_120) using torchao Float8Linear and optional TileKernels fused MoE routing.
1B parameters (996M), hidden_size=2048
Byte-level vocab: 272 tokens (256 UTF-8 bytes + 16 specials: BOS=257, EOS=258, PAD=256)
6 recurrent layers with MoE (32 experts, top-2 routing) + MLA attention
2 prelude + 2 coda dense… See the full description on the dataset page:
https://huggingface.co/datasets/CLIWorks/Spider-FLEXITOKENS-FP8.