Preference pairs for Direct Preference Optimization fine-tuning of
jspaulsen/halluci-mate-v1b,
a Qwen3-0.6B chess LLM trained from scratch on UCI moves.
Source: ~11,000 games of jspaulsen/halluci-mate-v1b vs. Stockfish (skill 5,
depth 12), evaluated with the halluci-mate
eval harness (scripts/eval.py vs-stockfish ... --sf-analyze).
Flavor: quality. The model's actual move is rejected; Stockfish's best move
on the same position (by… See the full description on the dataset page:
https://huggingface.co/datasets/jspaulsen/halluci-mate-v1b-dpo.