← The frontier
Technology & AIJun 28, 2026

PHF: Privileged Hidden Flow for On-Policy Self-Distillation

On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions.

On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions. Existing OPSD objectives supervise only the output distribution, so privileged context affects…

The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.