16,392 examples. Same as cot-oracle-truthfulqa-hint-cleaned but WITHOUT the strict clean_hint_answer_rate <= 5% filter.
Filtering: rollout match (model_answer == hint_answer for hint_used), diff >= 15pp, CI-based labeling, 50/50 resisted sub-balancing. But no per-rollout ground truth filter, so some hint_used examples may be coincidental (model would have picked that answer anyway).
See… See the full description on the dataset page:
https://huggingface.co/datasets/ceselder/TRUTHFULQA_OLD_BUT_KINDA_GOOD.