Views
No views yet
cap_20260716_115414,
6107 frames @ 10Hz), curated from CVAT task 8 annotations.person, bicycle, motorcycle, car, bus, truck.SAM3Image state_dict (1132 keys, incl. backbone +
trained segmentation_head.*), not nested under a "detector."
prefix like Meta's official release checkpoints — load with
checkpoint_path=None then model.load_state_dict(state_dict, strict=False)
manually, not via build_sam3_image_model's own checkpoint_path= loader.Sam3LossWrapper, matcher = BinaryHungarianMatcherV2 + o2m
BinaryOneToManyMatcher): Boxes loss_bbox=5.0/loss_giou=2.0; IABCEMdetr
loss_ce=20.0/presence_loss=20.0/pos_weight=10/focal α=0.25,γ=2; Masks
loss_mask=200.0/loss_dice=10.0, point-sampled (12544 points,
oversample_ratio=3.0, importance_sample_ratio=0.75), final decoder stage only.