Supervised fine-tuning data for heads-up No-Limit Texas Hold'em, 200 big
blinds deep. Each row is a single decision point: a natural-language
description of the game state, paired with the game-theory-optimal action
GTO Wizard chose in that spot.
Intended for instruction-tuning a chat LLM to play HU 200BB poker (see the
pokerbench agent it was built for).
Column
Description… See the full description on the dataset page:
https://huggingface.co/datasets/AYipppp/gtow-llama-sft-v3.