← The frontier
Technology & AISep 3, 2026

Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking.

We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.