The CoT module implements a 4-stage reasoning pipeline inspired by Alpamayo-R1 and AgentThink:
Scene Narration — Transformer decoder extracts 64 actor tokens and 32 road tokens from BEV, predicting class, distance, velocity, and initial threat per actor.
Risk Assessment — Per-actor risk analysis with self-attention (actors reason about interactions). Outputs TTC, collision probability, risk level (none/low/medium/high/critical), and identifies worst-case actor.
Causal Reasoning — 4-step autoregressive chain with causal masking:
Step 1: Situation assessment (what's happening)
Step 2: Hazard identification (what's dangerous)
Step 3: Action justification (why act this way)
Step 4: Action decision (what to do)
Safety Decision Gate — Monotonic safety constraint: the CoT can only make driving more conservative (reduce speed, increase braking), never more aggressive. Blends planner output with CoT override based on urgency × confidence.
Sensor Configuration
Default: 20 ultrasonic + 6 cameras at 20 mph
Cameras (6)
Name
Position
FOV
Resolution
cam_front_left
Front-left corner
120°
640×480
cam_front_right
Front-right corner
120°
640×480
cam_rear_left
Rear-left corner
120°
640×480
cam_rear_right
Rear-right corner
120°
640×480
cam_left_mirror
Left rearview mirror
90°
640×480
cam_right_mirror
Right rearview mirror
90°
640×480
Ultrasonics (20)
7 front bumper (spanning full width, angled -30° to +30°)