Model Card for Asymmetric-Executor-Diamond-Thinking
Model Summary
This model is a fine-tuned version of Qwen3-VL-8B-Thinking, designed for the Diamond Domain. It leverages the "Thinking" backbone's enhanced reasoning capabilities to handle the complex causal discovery required in the Diamond environment (e.g., inferring which button controls which wall based on observation history).
Training: Fine-tuned with LoRA on a curriculum of progressively harder reasoning steps (Q1 $\to$ Q5).
Intended Use
Best suited for deployment in the Diamond Domain or similar logic-heavy grid worlds where the agent must perform self-correction and verify complex state transitions (e.g., "The wall disappeared, so the button press was successful").
Performance
Diamond Domain Success Rate: ~94% (State-of-the-Art for this architecture).
Task Progress: Achieves 98% task progress, indicating high robustness in completing intermediate subgoals even in failure cases.