mask 6.2 % of framearea error ±1 %
Octopus segmenter
on held-out video
Pixel-accurate silhouette of a soft-bodied, camouflaging animal — with no prompt at inference. It scores above the 0.374 its own teacher architecture reaches zero-shot per frame, at a fraction of the size. Body area, the channel posture and masked-motion actually read, is accurate to about one percent.