Views
No views yet
[-1024, -512, -256, -128, -64, -32, -16, 0] env-steps; each past observation slot receives a learned per-slot soft gate weight, feeding a 2-layer MemoryTransformer that conditions the DiT flow-matching action head. A second prediction head performs action-conditioned forward prediction as an auxiliary loss. Trainable components: DiT + projector + moment tokens + top-4 LLM layers + gate + prediction head; the remainder of the VLM backbone is frozen.| Task Category | AAM (this ckpt) | HAMLET baseline |
|---|---|---|
| Counting | 5.0 | — |
| Permanence | 20.0 | ~19.5 |
| Reference | 12.5 | — |
| Imitation | 5.0 | — |
| Total avg | 10.6 | 16.5 |
nvidia/GR00T-N1.6-3B