Tracking a single pixel through a video sounds simple. It is one of the hardest problems in computer vision.
TAPNet, from Google DeepMind, tackles “Tracking Any Point” (TAP): given any query point on any frame, predict where that exact physical surface point sits in every other frame, and whether it is visible or occluded.
Why this is different from classic approaches:
- Optical flow only links adjacent frames, so small errors accumulate into drift over long sequences
- Object trackers follow boxes or masks, not specific surface points
- Keypoint detectors only work on predefined features like joints or corners
How TAPNet works:
- Each frame is encoded into a dense feature grid by a convolutional backbone
- The query point’s feature is compared against every location in every frame, building a cost volume
- A small network reads that cost volume to output a position and an occlusion probability per frame
- Training relies heavily on Kubric, a synthetic renderer where ground-truth mo...
Suggested Credits
Tags, Events, and Projects