Technology & AIJul 30, 2026
$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation
On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort.
On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is…
Sign in to learn & save →
The frontier is open to all. Sign in to learn this from first principles and save it to your knowledge base.