# @stem_antics on Instagram

- **Type:** Video
- **Original URL:** https://www.instagram.com/p/DeJ7jeNAVg6
- **Gondola URL:** https://gondola.cc/posts/71427941-stem-antics-instagram
- **Thumbnail:** https://img.gondola.cc/tr:w-,h-,fo-auto/postThumbnails/8d22de27f7.jpg
- **Posted:** 2026-10-06T13:50:52.000+00:00
- **Account Owner:** Stem Antics (@stem_antics) — https://gondola.cc/stem_antics

## Caption

Tracking a single pixel through a video sounds simple. It is one of the hardest problems in computer vision.

TAPNet, from Google DeepMind, tackles “Tracking Any Point” (TAP): given any query point on any frame, predict where that exact physical surface point sits in every other frame, and whether it is visible or occluded.

Why this is different from classic approaches:
- Optical flow only links adjacent frames, so small errors accumulate into drift over long sequences
- Object trackers follow boxes or masks, not specific surface points
- Keypoint detectors only work on predefined features like joints or corners

How TAPNet works:
- Each frame is encoded into a dense feature grid by a convolutional backbone
- The query point’s feature is compared against every location in every frame, building a cost volume
- A small network reads that cost volume to output a position and an occlusion probability per frame
- Training relies heavily on Kubric, a synthetic renderer where ground-truth motion of every pixel is known exactly

The team also released TAP-Vid, a benchmark built from real videos (Kinetics, DAVIS) plus robotic and synthetic scenes, with points annotated by humans. Performance is scored with Average Jaccard, which combines position accuracy across multiple pixel thresholds with correct occlusion prediction.

The follow-up, TAPIR, added a temporal refinement stage and delivered a large jump in accuracy, and later work like BootsTAP used millions of unlabeled real videos to keep improving.

Why it matters: precise long-range point tracks give robots, animators, and scientists a dense record of how things actually move.

#stemantics #computervision #deeplearning #machinelearning #robotics

## Stats

- **Views:** 3,771
- **Likes:** 18
- **Shares:** 0
- **Comments:** 2

## Tags

robotics, deeplearning, machinelearning, computervision, stemantics

---
Copyright (c) Gondola