Precomputed per-word predictors for computational models of human sentence
processing: word surprisal, several kinds of entropy, and unigram surprisal,
scored over reading corpora (eye tracking, self-paced reading, N400, maze).
Every row says what text unit it describes and what unit its number is in:
provo
gpt2_ft_entropy_provo
token_entropy
provo_1
4
Apple
8.4213
bits… See the full description on the dataset page:
https://huggingface.co/datasets/samuki-hf/psycholing-predictors.