Views
No views yet
step_0100000.pt (100K gradient steps)| Metric | Final Value |
|---|---|
| World Model ELBO | 22.44 |
| Reconstruction Loss | 0.001 |
| KL Divergence | 19.20 |
| Reward Prediction | 3.17 |
| Actor Return (mean) | 124.33 |
| Critic Loss | 2.50 |
| Training Time | ~2.6 hours |
| Hardware | NVIDIA DGX Spark (GB10 GPU, 128 GB unified memory) |
| Metric | h=1 | h=5 | h=10 | h=13 | h=25 |
|---|---|---|---|---|---|
| RMSE | 5.001 | 5.167 | 5.286 | 5.283 | 5.243 |
| MAE | 2.130 | 2.168 | 2.207 | 2.205 | 2.197 |
| WMAPE (%) | 71.7 | 72.3 | 72.6 | 72.6 | 72.4 |
| NDR(h) | 1.00 | 1.033 | 1.057 | 1.056 | 1.049 |
| Method | Mean Return | IQM |
|---|---|---|
| Cost-plus (25%) | 54.8 | 54.8 |
| Static XGBoost | 87.2 | 85.6 |
| Competitive Matching | 42.1 | 41.5 |
| DQN | 68.9 | 65.2 |
| PPO | 76.4 | 72.8 |
| SAC | 82.3 | 79.6 |
| DreamPrice | 124.3 | 117.4 |
| Component | Specification |
|---|---|
| Backbone | Mamba-2 SSM (d_model=512) with GRU fallback |
| Stochastic latent | 32 categorical variables x 32 classes (z_dim=1024) |
| Posterior | DRAMA-style decoupled: q(z_t | x_t) |
| Observation decoder | 3-layer MLP: cat(h_t, z_t) -> obs_dim |
| Demand decoder | Causal: theta * log(price) + MLP(z_t, store_features) |
| Reward ensemble | 5 independent heads, twohot distributional (255 bins) |
| Continue head | Linear -> sigmoid |
| MOPO pessimism | r_pessimistic = r_mean - lambda_lcb * r_std |
1@article{sathish2026dreamprice,
2 title = {DreamPrice: A Learned World Model for Retail Pricing via Mamba-2 Recurrence and Causal Demand Identification},
3 author = {Sathish, Sharath},
4 year = {2026},
5 url = {https://github.com/SharathSPhD/dreamprice}
6}