Views
No views yet
| Metric | Value |
|---|---|
| vs Champion (V7 Human BC) | 58.5% |
| vs V1 (83% Rule) | 72.0% |
| vs Rule Agent | 95.0% |
| State Encoding | 716-dim (714 + M-value + Power) |
| Architecture | LSTM(2-layer) + ResNet MLP |
| Training Steps | 3,810,000 / 5,000,000 |
| Training Method | Deep Monte Carlo (DMC) with N-step TD |