This dataset contains synthetic short murder mysteries for training and testing
the tiny Clue 250K language model. The examples use a fixed set of names,
locations, weapons, and wound descriptions. Each mystery asks the model to infer
the murderer, or answer Unknown when the clues do not identify exactly one
person.
The companion code and trained demo model live in the GitHub repository for
Clue 250K. The dataset is… See the full description on the dataset page:
https://huggingface.co/datasets/gszauer/Clue250K.