Pre-classification-layer representations of the 10,000 harmful construction
points from entfane/construction_points,
extracted from three author-trained guardrail classifiers of Beyond
Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
(arXiv:2605.10901).
These are the exact cached activations used to reproduce the zero-knowledge
guardrail certificates in the zkrobustness/zk prototype (the deterministic
hyper-rectangle SAT/UNSAT… See the full description on the dataset page:
https://huggingface.co/datasets/paberr/guardrail-head-embeddings.