Training data for Natural Language Autoencoder adapters: a diverse text
corpus paired with token-prediction-style descriptions of what a model is
computing at each network depth. Used to train the activation verbalizer (AV)
and reconstructor (AR) in the
nla-at-home project — a DIY
replication of Anthropic's
Natural Language Autoencoders.
5,213 source texts across 55 categories
(code, math, grief, dharma, medical, multilingual… See the full description on the dataset page:
https://huggingface.co/datasets/anicka/nla-at-home-corpus.