Companion to tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.
For every one of the 153,796 5-second EEG windows in that dataset, this repo provides a
frozen text embedding of the window's transcript — the CLIP target for EEG→text decoding,
the text analogue of the w2v (wav2vec2-XLSR-53) speech target in the main dataset.
Built to extend the EEG→speech replication of Sato et al. 2024, "Scaling… See the full description on the dataset page:
https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-qwen3-text.