LLM reasoning traces for BTC/USD trading decisions. Each row is one hourly
trading decision: the model receives account state + multi-timeframe OHLCV
market data, reasons step-by-step in
tags, and outputs a discrete
action (N). The dataset stores the full chain-of-thought
alongside the resulting reward and next state — making it suitable for
offline RL, imitation learning, or reward modelling.