Private dataset of on-policy model-organism transcripts labelled
honest/deceptive, for lie-detection research.
Do not redistribute.
model — HuggingFace repo id of the model organism that generated the transcript.
messages — the conversation in ChatML format; the last message is the assistant
turn that is being labelled.
deceptive — bool; whether the last assistant message… See the full description on the dataset page:
https://huggingface.co/datasets/AlignmentResearch/hidden-goal-model-organism-deception-dataset-nemotron3-super-v1.