This workspace is a runnable benchmark and distillation scaffold for food-and-feeding perception in mukbang / human-feeding video. The near-term teacher stack is:
Sapiens2 INT4-G128 for human/body segmentation, normals, pointmap/depth proxy, and matting.
SAM3.1 NVFP4 for open-vocabulary food instance segmentation and coarse food naming.
Sapiens2 segmentation as a hard exclusion mask so food masks do not claim pixels already labeled… See the full description on the dataset page:
https://huggingface.co/datasets/Reza2kn/stallion-backup-nomnomlabel-20260713.