Training corpus for the monitoring experiments of the Stream-LLM models
(Stream-Qwen3.5-27B,
Stream-Qwen3-8B).
Each sample is a ten-column grid where every column is one cognitive
channel; per row, each channel contributes one short phrase (or silence -).
raw
raw/train.parquet
3874
Original machine-generated grids in natural language.
processed
processed/train.parquet3864
Tokenized with the Qwen3.5-27B… See the full description on the dataset page:
https://huggingface.co/datasets/JonasGeiping/stream-data.