Views
No views yet
pile Quadratic/bilinear attention causal language model trained with the tensor-mars research stack. This repository packages the final checkpoint, configuration, and reference model code. ## Training configuration ```yaml batch_size: 384
## Metrics - **train_loss**: 3.9820… See the full description on the dataset page: https://huggingface.co/datasets/melephant/2l-bilinear-attn.