Using task-specific CLIP text prompts injected into cross-attention layers, SILICA establishes a spatial hierarchy that resolves foreground-background visual ambiguities inherent to transparent surfaces.
This work has been accepted for publication at IROS 2026.
For setup instructions, Python standalone inference, environment setup via
uv/
pip, and ROS2 (Humble) integration,
please refer directly to the Official GitHub Repository.